Skip to main content
Ask about what is on your screen and the assistant answers with that context. It only looks when you let it. Set it up in Settings → Assistant → Screen vision. You need a vision-capable model. Most curated local models qualify (the catalog marks them “Vision”), as do the mainstream cloud multimodal models.

Screen access

Every screenshot that goes to the model is shown in the conversation, in both active modes, so there is always a visible record of when the assistant looked.
macOS. Screen vision needs macOS Screen Recording permission. Grant it in System Settings → Privacy & Security → Screen Recording, then restart SpeakoFlow for it to take effect.

What actually gets sent

SpeakoFlow captures the monitor under your mouse cursor, so on a multi-monitor desk you get the screen you are actually working on. Only a compact thumbnail is stored in the conversation and in History. The full-resolution frame is sent to the model once and never written to disk. Thumbnails show inline in the panel and can be clicked to enlarge.
When I ask, When the message sends. Default: When I ask. Shown in Manual mode only.This changes the timing for voice questions only, where there is a real gap between starting to ask and finishing.
  • When I ask (default). Grabs the frame the moment you press the hotkey, so it captures what you were looking at when you started talking, not whatever is on screen after you finish.
  • When the message sends. Grabs it after you stop talking and the speech is transcribed.
Typed messages always capture on send, since the panel is already in front of you either way.
The frame is JPEG-compressed down to a budget picked from your provider, because “send the sharpest possible image” and “the request succeeds” are not the same goal.
The panel says so and names the problem rather than silently dropping the screenshot. Switch to a vision-capable model, or ask again with screen vision off.
Last modified on August 7, 2026