Note
October 8, 2026
2 min read
Does Bob know I can’t see the screen?
By Cristiano Pierry
Voice conversations work better when an agent knows whether its user can see the screen or needs answers that stand on their own.

I was talking to Bob, my AI agent, while driving this morning. He had put a draft in our chat and was creating an illustration. Both were useful. I needed him to read one and describe the other.
It made me wonder what he knows about how I’m talking to him. When I use voice inside the app, I can talk and look at the screen together. When I call my dot while driving, the conversation needs to work entirely through audio. In this conversation, Bob said he wasn’t given a reliable indication of which interface I was using or whether I could see the screen.
I’d like dots to receive that context. A call could default to spoken answers that stand on their own, with descriptions of images and no assumption that I can inspect a link or read something in chat. In-app voice could make use of the screen when I’m able to look at it.
The interface alone won’t tell the whole story. I might start voice in the app and then put my phone in my pocket. A simple “I’m listening only” setting, or telling Bob that directly, should override the default.
We spend a lot of time improving what these systems can do. Knowing whether I can see what they’re showing me seems like a useful part of doing it well.
This writing reflects my personal perspectives on product management, AI, and content discovery. It does not represent the official position of my employer or any affiliated organization.