Everything SpeakoFlow can do.
SpeakoFlow turns your voice into text, right where you're working. Press a hotkey and talk, and your words are typed into whatever app you're using. This guide covers every feature, setting, and shortcut, on Windows, macOS, and Linux.
Speech-to-text runs locally on your machine, so your voice never leaves your device. The AI assistant runs on any model you choose, from a fully offline built-in model to your own local server or a cloud provider with your own key. You decide how much stays on your machine.
Talk, it types
Words land in any app, live as you speak or all at once when you stop.
Say "Hey Flow"
It writes the reply, email, or draft and pastes it for you.
A voice you can ask
A floating panel that answers by voice or text, and can read replies aloud.
On your device
No account, no telemetry. Optional features stay off until you turn them on.
How to read this guide
Each section is self-contained, so you can jump straight to what you need from the sidebar. New to the app? Start with Getting started and Keyboard shortcuts — that's genuinely all you need for day one. Everything else is here when you want to go deeper.
Throughout the docs you'll see platform tags like macOS for steps that differ by operating system, and Experimental for features that are still being refined.
Getting started
Install the app, pick a model, and dictate your first sentence. The whole thing takes a couple of minutes.
1. Install
Download the latest build for your platform from the Releases page. SpeakoFlow runs on Windows, macOS, and Linux.
2. Run the setup wizard
On first launch, a short wizard walks you through two steps:
- Hear you. Choose a speech-to-text model. This turns your voice into text right on your machine. Pick one now — you can switch anytime in Settings. A fast English model or a real-time multilingual model are good starting points.
- Give it a brain (optional). Optionally download a small local model for the assistant and AI cleanup. You can skip this and set it up later under Settings → Assistant → On my device.
Models download in the background. You can start using the app right away, and any other model is one click away in Settings.
3. Grant permissions
SpeakoFlow needs two permissions to work:
- Microphone access — required to hear your voice for transcription.
- Accessibility access macOS — required to type transcribed text into your applications.
If you skip these during setup, you'll be prompted again the first time they're needed. See Troubleshooting if a permission gets stuck.
4. Say your first sentence
Open any text field — an email, an editor, a chat box — then hold your Dictate shortcut and talk. Let go, and your words type themselves. That's the whole loop.
| Action | Windows | macOS | Linux |
|---|---|---|---|
| Dictate | Left Ctrl+Left Win | Option+Space | Ctrl+Space |
| Ask the assistant | Left Ctrl+Left Alt | Option+Ctrl+Space | Ctrl+Alt+Space |
Every shortcut is rebindable in Settings → General → Shortcuts. The full list is in Keyboard shortcuts.
How SpeakoFlow works
A quick mental model that makes the rest of the docs click into place.
One hotkey layer over your real work
SpeakoFlow sits quietly in your system tray and does nothing until you press a shortcut. It's a layer over whatever app you're already in — you never switch to it to dictate. Three jobs live under that hotkey layer:
- Dictation — local speech-to-text typed into the active app. Your audio never leaves the machine.
- Generate with Flow — start a dictation with an activation phrase and the AI writes the finished result.
- The assistant — a floating panel you summon to ask questions by voice or text, optionally with screen vision and spoken answers.
The dictation pipeline
When you dictate, your audio flows through a local pipeline: your microphone → voice-activity detection (which filters out silence) → a speech-to-text model running on your GPU or CPU → text inserted into the active app via the clipboard or direct typing. Nothing in that chain touches the network.
Local-first, your choice
Speech-to-text is always local. The assistant is where you choose your comfort level:
Built-in (offline)
Download a small local model and run it fully on your machine. No key, no network.
Local server
Point SpeakoFlow at Ollama or LM Studio running on your own machine.
Cloud
Bring your own API key for any OpenAI-compatible provider.
Settings, organized the same way
The Settings window mirrors this guide: General (shortcuts, microphone, appearance), Dictation (the model and how text is cleaned up), Assistant (brain, voice, screen vision, web search), plus Profiles, Memory, History, and Debug.
Keyboard shortcuts
Every shortcut is configurable in Settings → General → Shortcuts. These are the defaults out of the box.
| Action | Windows | macOS | Linux |
|---|---|---|---|
| Dictate Start/stop recording, type it out | Left Ctrl+Left Win | Option+Space | Ctrl+Space |
| Dictate & clean up Dictate, then AI-clean before pasting | Ctrl+Shift+Space | Option+Shift+Space | Ctrl+Shift+Space |
| Ask the assistant Voice question into the panel | Left Ctrl+Left Alt | Option+Ctrl+Space | Ctrl+Alt+Space |
| Show / hide panel Toggle the floating assistant | Ctrl+Shift+A | Option+Ctrl+A | Ctrl+Alt+A |
| Cancel Stop a recording or streaming reply | Unset by default — bind any key (e.g. Esc) in Settings | ||
| Debug mode Reveal diagnostics | Ctrl+Shift+D | Cmd+Shift+D | Ctrl+Shift+D |
Why the modifier-only defaults on Windows? Holding Left Ctrl+Left Win keeps every letter and the Space bar free, so the shortcut can't collide with text shortcuts inside your apps. Left-side keys specifically, so an international AltGr layout won't trigger the assistant by accident.
Recording behavior: Hold vs. Tap
In Settings → General you choose how the Dictate shortcut behaves:
- Hold (default) — records while the shortcut is held down, and types out when you let go.
- Tap — one press starts recording, the next press stops it. No holding.
For hands-free recording without switching modes, see Hands-free & tap-to-lock below.
Dictation basics
Press a hotkey and talk. Your words type into any app — email, editor, chat, anywhere you can put a cursor.
The loop
- Put your cursor where you want text to appear.
- Hold the Dictate shortcut and speak naturally.
- Release. SpeakoFlow transcribes and inserts your words into the active app.
Voice-activity detection trims silence automatically, so brief pauses while you think won't add gaps. If you prefer tap-to-start over hold, switch Recording behavior to Tap in General settings.
Talking is roughly 3× faster than typing — around 150 words a minute spoken versus 45 typed. The speed adds up fast across a day of messages and notes.
Hands-free & tap-to-lock
Lock recording so you can let go of the keys and keep talking — useful for long dictation or when your hands leave the keyboard.
While you're holding your record shortcut, tap the tap-to-lock key (Space by default) once. Recording locks hands-free, so you can release the shortcut and keep talking. Press the shortcut again to stop and type it out.
Pick any key that isn't part of your record shortcut, or clear it to turn the feature off. The assistant shortcut has its own separate tap-to-lock key.
The overlay shows "Recording hands-free — press the hotkey again to stop" once you're locked, so you always know which mode you're in.
Live vs. batch transcription
Choose whether words appear as you speak, or all at once when you stop.
By default, SpeakoFlow transcribes in a single batch when you finish — accurate and simple. With live transcription on, text streams in as you talk (on models that support streaming).
Transcribe speech as you talk, instead of all at once when you stop.
While live transcription runs, show a larger overlay card with the running text instead of the compact pill. Requires live transcription.
Streaming-capable models (like the Moonshine V2 and Nemotron streaming models) are the best fit for live mode. See Transcription models for which ones stream.
The recording overlay
A small on-screen indicator that shows SpeakoFlow is listening, so you're never guessing whether it heard you.
Overlay style
Set in Settings → General → Overlay. Controls how the recording overlay looks while you dictate:
- None — hides the overlay entirely.
- Minimal — a compact pill that shows recording state.
- Live — a readable card that shows your words as you speak (for models that support it).
Overlay position
Where the overlay sits on your screen: Bottom (default), Top, or None.
The overlay also surfaces status while you work — Getting mic ready, Recording, Transcribing, Writing with Flow, or Looking at your screen — and lets you cancel in progress.
Linux On native GNOME/Wayland the overlay can't always stay on top of other windows. SpeakoFlow works around this automatically — see Troubleshooting.
Translate to English
Speak another language and get clean English out — on your device, with a capable model.
Translate speech from other languages to English as it's transcribed. Availability depends on the model — translation-capable models include Whisper, Canary, and Cohere. If your current model doesn't support it, the setting notes that and you can switch models.
Custom words
Teach SpeakoFlow to spell names, products, and phrases your way.
Add names, terms, or short phrases that the model tends to get wrong, and SpeakoFlow will correct them in your transcripts. For example, add Kokoro if it keeps writing "Coco row."
Custom words shine for non-Western names and specialized vocabulary — great for names from Nepali, Indian, and other languages that generic models spell phonetically. Pair them with a strong multilingual model like Whisper for the best results.
The sensitivity of custom-word correction can be tuned with the Word correction threshold in Debug settings.
Text replacements
Swap a phrase you say for text you use often — like turning "my email" into your actual address.
Create rules that map a spoken phrase to inserted text. Say "my email" and SpeakoFlow inserts alex@example.com. Rules are managed in Settings → Dictation (Advanced → Text replacements) and can be imported and exported as a list.
Each rule has
- Spoken phrase — what SpeakoFlow hears (e.g. "my email").
- Insert this text — what it types instead.
- Enabled — toggle a rule on or off without deleting it.
Advanced options
- Regex — match with a regular expression instead of a literal phrase.
- Trim before / Trim after — remove surrounding whitespace around the match.
- Caps — force the inserted text to None, UPPERCASE, lowercase, or Capitalize.
Magic tokens. Use [date], [time], [uppercase], [lowercase], [capitalize], or [nospace] inside the inserted text for dynamic values and formatting.
Spoken emoji
Say the name of an emoji and get the emoji.
Say "happy emoji" or "thumbs up emoji" to insert 😊 or 👍. Common aliases and minor speech-to-text errors are handled on your device; unknown names stay as plain text.
Output & pasting
Fine control over exactly how transcribed text is inserted into the active app. Most people never need to touch these — the defaults just work.
How transcribed text is inserted: Clipboard (paste via Ctrl/Cmd+V, with Ctrl+Shift+V and Shift+Insert variants), Direct (typed key by key), None, or an External script you point to.
The Linux tool used for the Direct paste method. Leave on Auto (recommended) unless you have a reason to change it.
Whether the transcription stays in your clipboard after pasting: Don't modify clipboard (restores what you had) or Copy to clipboard (leaves the transcript there).
Send a key combination automatically after text is inserted — Off, Enter, Cmd+Enter, Super+Enter, or Ctrl+Enter. Handy for firing off chat messages hands-free.
Add a space after inserted text so consecutive dictations don't run together.
If text is occasionally cut off or the wrong text pastes, the Paste delay and Extra recording buffer in Debug settings can help.
Transcription models
The model is what turns your voice into text. Every model runs on your machine — they trade accuracy for speed, so pick the one that fits your hardware and languages. Manage them in Settings → Dictation and the model catalog.
Current models
| Model | Best for | Notes |
|---|---|---|
| Parakeet (V2 / V3 / unified) | Fast everyday dictation | V2 English-only; V3 covers 25 European languages. Blazingly fast. |
| Nemotron streaming | Real-time multilingual | Streams as you speak. 28 languages. Streaming |
| Canary (180M Flash / 1B v2) | Speed + translation | 180M is tiny and instant; 1B covers 25 European languages. Both translate to English. |
| Whisper (Small / Medium / Turbo / Large) | Broad multilingual accuracy | Medium supports nearly 100 languages. The most versatile family. |
| Moonshine (Base / V2 Tiny·Small·Medium) | Ultra-fast English | English only; handles accents well. V2 variants stream. Streaming |
| Cohere | Highest accuracy | Large and slower, but the most accurate. Many languages. |
| SenseVoice | East Asian languages | Very fast. Chinese, English, Japanese, Korean, and Cantonese. |
| GigaAM v3 | Russian | Russian only. Fast and accurate. |
| Breeze ASR | Taiwanese Mandarin | Tuned for Taiwanese Mandarin with English code-switching. |
Older models run on the previous engine and are grouped separately in the catalog. They're still fully supported, but most people won't need them.
Managing models
- Browse Downloaded models and Available to download in the catalog.
- Filter by All, Multi-language, Translation, or search by name.
- Downloads resume automatically if interrupted, and retry on flaky connections.
- Delete a model to free disk space; you'll be warned if it's your active model.
Languages & accents
Getting names and non-English speech right comes down to the model you pick — and a couple of settings that reinforce it.
The language for speech recognition. Setting a specific language (instead of Auto) can improve accuracy noticeably. Some models detect the language automatically and don't expose this setting.
Which model for which languages
- Whisper Medium/Large — the widest coverage, nearly 100 languages. The best all-rounder for mixed or non-Western speech.
- Nemotron streaming — 28 languages in real time.
- Parakeet V3 / Canary 1B — 25 European languages.
- SenseVoice — Chinese, English, Japanese, Korean, Cantonese.
Non-Western names. Whisper models are particularly strong at names and words that trip up English-only models — for example Nepali, Indian, and other South Asian names that don't follow standard English spelling. If you dictate names like these often, use a Whisper model and add the tricky ones to Custom words so they're spelled your way every time.
Acceleration & GPU
SpeakoFlow uses your GPU when it can, falling back to CPU otherwise. These live under Settings → Dictation → Advanced.
Auto uses the GPU when available. Metal on macOS, Vulkan on Windows, OpenBLAS + Vulkan on Linux.
For Parakeet, Canary, and Moonshine models. DirectML is experimental.
Auto picks the dedicated GPU. Choose a specific device if you have more than one.
Custom models
Beyond the built-in catalog, you can pull GGUF models straight from Hugging Face.
In the model catalog, choose Search Hugging Face to find a GGUF model, pick a version, and download it directly to your device. This is an advanced option — hardware needs and compatibility can vary, and custom models aren't officially supported.
- Search by model or creator, or browse popular GGUF models.
- Choose a quantization — Q4_K_M is usually the best balance of quality, speed, and memory.
- For assistant models, optionally include screen vision, which downloads an extra file so the model can understand images.
Memory & the microphone
Settings that trade a little memory or battery for speed and reliability.
Free memory when the model has been idle for a while. Options range from Never and Immediately to after 2 / 5 / 10 / 15 minutes or 1 hour. The model reloads automatically the next time you dictate.
Keep the mic warm so the first words are never clipped. Uses a little more battery.
Faster back-to-back dictation by leaving the audio stream open. May affect Bluetooth audio quality.
Generate with Flow
Begin a dictation with an activation phrase and the AI writes the finished result — a reply, an email, a draft — and pastes it for you.
How to use it
Turn on Generate with Flow in Settings → Dictation, then dictate a command that starts with your activation phrase:
"Hey Flow, write a polite email asking to reschedule Friday's call to Monday."
Flow generates the finished email and pastes it where you're typing. Anywhere the phrase isn't at the start, the same words paste as normal dictation — so Flow never gets in the way of regular use.
Flow activates only when a dictation begins with this phrase. Rename it to anything you like; leave it empty to reset back to "Hey Flow". Flow uses your assistant's model to generate.
Let Flow take one screenshot when a command clearly refers to your screen — like "look at this email and draft a reply." The model decides whether it's needed. This is separate from the assistant's screen access.
Flow needs an assistant model configured. If none is set up, your words are pasted as ordinary dictation and the overlay lets you know. Set one up in the assistant's brain settings.
AI cleanup
Experimental Speak messy, paste polished. Cleanup runs a language model over your dictation to strip filler, fix grammar and punctuation, and optionally match a tone.
AI cleanup is still in development and off by default. It adds a short delay after you speak, since the model runs after transcription. If it fails or times out, your raw text is pasted instead — you never lose your words.
How to use it
Cleanup applies to dictation only (not the assistant chat) and runs on its own shortcut — Dictate & clean up. So you can dictate raw with your normal shortcut, and dictate-then-clean with the other whenever you want a polished result. Turn it on and tune it in Settings → Dictation → AI cleanup.
How much the AI rewrites your words:
- Light — fixes only mistakes and filler.
- Balanced — also trims repeated points and false starts.
- Aggressive — tightens rambling into concise, well-structured text.
All levels keep your facts and meaning intact.
When a word is clearly wrong for the sentence — a homophone or a near-miss pronunciation — replace it with the word you meant. Helpful for accents; the AI only corrects when the intended word is obvious.
Choose how the cleaned dictation should sound: None (cleanup only), Formal, Casual, Professional, Friendly, or Concise. "None" fixes the text without intentionally changing its voice. You can also create your own — see Writing styles.
By default cleanup uses your assistant's model. You can also run it privately on a downloaded local model, on a cloud provider, or on Apple Intelligence macOS (fully on-device, no key — requires an Apple Silicon Mac on macOS Tahoe 26.0+ with Apple Intelligence enabled).
Pick a template for cleaning up transcriptions or write your own. Your transcribed text is sent to the model automatically — just write instructions, e.g. "Fix grammar and punctuation, then rewrite as a polished, professional message."
How long to wait for the model before giving up and pasting your raw transcription instead. On the built-in local model, allow enough time for it to load on first use.
Writing styles
Beyond the built-in tone presets, you can save your own rules for exactly how cleaned dictation should read.
In the AI cleanup Style preset section, choose Create style and give it:
- Style name — e.g. "Calm and clean."
- Style instructions — describe precisely how the wording should change. For example: "Remove profanity and replace it with calm, neutral wording."
These rules shape the wording after cleanup and stay separate from your cleanup prompt, so you can mix any prompt with any style. Edit, update, or delete your saved styles anytime.
Assistant panel
A floating chat you summon with a hotkey. Ask by voice or text, get streaming answers, and have them read back aloud — without leaving the app you're in.
Using the panel
- Ask by voice — hold Ask the assistant and talk; the answer streams into the panel.
- Ask by text — type in the message box.
- Show / hide — toggle the panel with its shortcut. It remembers the conversation while open.
- Collapse to a pill — shrink the panel to a small pill you click to talk, then expand again when you want the full view.
- Attach — add an image or file, or snip a region of your screen to ask about.
- Regenerate — re-ask for a fresh answer; the old one stays in History.
The panel also exposes quick toggles for screen vision, web search, and voice replies, and a persona switcher in its header (see Profiles).
Brain: local or cloud
The assistant runs on whatever model you choose. This is the single most important assistant setting — set it in Settings → Assistant → Assistant brain.
On my device
Download a private local model. It stays on this machine and needs no API key. First use downloads a small local engine automatically, one time, then runs fully offline.
Cloud provider
Answer through an OpenAI-compatible cloud provider using your own API key.
Choosing a local model
The catalog recommends three conversation-first choices: quickest, recommended, and more capable but slower. SpeakoFlow detects your GPU and reports graphics memory so you can gauge which size will run well.
| Tier | Example | Notes |
|---|---|---|
| Small / quick | Gemma 4 E2B, Gemma 3 1B | Runs on any machine. Great for tidying writing and simple questions. |
| Recommended | Gemma 4 E4B | A strong balance of quality and speed for conversation. |
| More capable | Gemma 4 12B, Qwen vision models | Sharper answers, some can see your screen — but slower, best with a strong GPU. |
Vision-capable models download an extra image file automatically so they can read your screen. You can also browse Hugging Face for other GGUF models — see Custom models.
Providers & keys
For cloud or local-server brains, SpeakoFlow speaks the OpenAI-compatible API, so it works with most providers.
Supported providers
OpenAI, Anthropic, Azure OpenAI, OpenRouter, any custom OpenAI-compatible endpoint, and local servers like Ollama or LM Studio. Provider quirks are handled for you — Anthropic's x-api-key, Azure's api-key and base-URL normalization, and OpenRouter's headers.
The provider that answers assistant questions.
The OpenAI-compatible base URL. Editable for Custom, Local, and Azure. Azure URLs are normalized automatically (e.g. to /openai/v1). For a local server, point it at your machine, e.g. http://localhost:11434/v1.
The key for the selected provider, stored securely in your OS keychain. Shared with AI cleanup, since both use the same chat client.
The model or deployment name to use for answers, e.g. gpt-5-mini. Load the provider's model list, or type a name directly.
Local model tuning
How much the local model can keep in mind at once — the system instructions, your chat history, any screenshot (a single image can eat a big chunk), and its reply. Bigger holds more and keeps screen vision and web search reliable, but uses more memory. 16384 is a good step up if your machine has the RAM; you can lower it later to save memory.
Free memory by unloading the local model after it's been idle this long. It reloads automatically on next use.
Sets the assistant's baseline behavior. Profiles can override this per-persona (see Profiles).
API keys live in your operating system's keychain, not in a plain settings file. The assistant only ever contacts the provider you choose — which can be a fully local one.
Behavior & memory
Tune how the assistant carries a conversation and how long its answers run.
How many recent messages are sent as context each turn. Higher keeps more of the thread in mind; 0 disables it and treats every message as standalone.
When a conversation grows past the model's context window, automatically condense older messages into a running summary so the chat keeps its context instead of forgetting the start. A subtle indicator shows while it runs.
Preferred reply length: No effect (leaves your prompt untouched), Short, Medium, or Long. Individual profiles can override this.
This is short-term, per-conversation context. For durable facts the assistant remembers across chats, see Personal memory — a separate, opt-in feature.
Screen vision
Ask about what's on your screen and the assistant answers with that context. It only looks when you let it.
Screen vision needs a vision-capable model (many Gemma and Qwen models qualify). Set it up in Settings → Assistant → Screen vision.
- Off — screen capture is disabled; camera and region controls are hidden.
- Manual — you choose when to attach the full screen or a region.
- Agent decides — the model may request a capture when it needs current visual context. Every screenshot sent is shown in the conversation.
For voice questions: capture the moment you start asking (When I ask) or wait until the message is sent (When the message sends). Typed messages always capture on send.
macOS Screen vision needs macOS Screen Recording permission. Grant it in System Settings → Privacy & Security → Screen Recording, then restart SpeakoFlow for it to take effect.
SpeakoFlow captures the monitor under your mouse cursor, so multi-monitor users get the screen they're actually working on. Only a compact thumbnail is stored in history; the full-resolution frame is sent once to the model and not kept.
Web search
Optional. When a question needs current, factual information, the assistant can look it up. The model decides when to search — you don't have to ask.
Enable it in Settings → Assistant → Web search. Off until you turn it on.
Let the assistant search the web when a question needs current information.
Two ways to search
If you use OpenRouter, turn this on to let OpenRouter run the search itself (its :online mode) using your OpenRouter credits — no separate search API key needed. When off, the app searches with your own search provider below.
Choose a provider for app-run search. Every provider has a free tier:
- Serper — Google results, fast. A good default.
- Brave Search
- Tavily — AI search
- Exa — neural search
- SerpAPI — Google
Key for the selected search provider (not needed when using OpenRouter's built-in search).
Quality & cost controls
Read the full text of each result instead of snippets. Better answers, but slower.
How many results to read.
Quick for simple facts, Standard for most questions, Deep for hard ones.
Let the local model decide when to search. Off uses a fast keyword check instead.
A daily cap on Firecrawl credits before search pauses. 0 means no limit.
Run a sample search with the current provider and key to confirm it's working before you rely on it.
Voice output (TTS)
Have the assistant read its replies aloud. Runs locally and free with Kokoro, or through a cloud voice with your own key.
Set up in Settings → Assistant → Voice output.
Read the assistant's replies out loud.
Stop any reply that's still playing when you start dictating, so you never talk over it.
Engines
| Engine | Runs | Needs |
|---|---|---|
| Kokoro Local | On your machine | One-time voice model download, then fully private. Free. |
| OpenAI-compatible | Cloud | Base URL, API key, model (e.g. gpt-4o-mini-tts), voice (e.g. alloy, coral, sage). |
| OpenRouter | Cloud | Your OpenRouter key. |
| ElevenLabs | Cloud | API key, model (e.g. eleven_v3, eleven_multilingual_v2, or eleven_flash_v2_5), voice ID. |
| Azure AI Speech | Cloud | Speech URL, key, voice (e.g. en-US-AvaMultilingualNeural). |
Picking an ElevenLabs model. eleven_v3 is the newest and most expressive, eleven_multilingual_v2 is the consistent-quality choice across 29 languages, and eleven_flash_v2_5 is the fastest (ultra-low latency, ~75 ms). Model names are the API IDs — type them exactly as shown.
Voice settings
The voice used for spoken replies. Kokoro has its own set; cloud engines let you load and pick from their available voices. Examples:
- OpenAI-compatible —
alloy,ash,ballad,coral,echo,fable,nova,onyx,sage,shimmer,verse. - ElevenLabs — a voice ID from the Voices tab in your ElevenLabs dashboard.
- Azure — a neural voice like
en-US-AvaMultilingualNeural, or an HD voice such asen-US-Ava:DragonHDLatestNeural. Click Load voices to browse.
fp32 is highest quality; q8 is faster on CPU. Changing this reloads the model.
1× is normal. Adjust faster or slower to taste. (ElevenLabs supports 0.7×–1.2× only.)
Play a short sample with the current settings before you commit.
Panel appearance
Make the floating panel fit your desktop. A live preview shows your changes as you make them.
Size of the chat text in the assistant panel: Small, Medium, or Large.
Overall size of the floating panel: Small, Compact, Standard, or Large. You can also drag its edges to resize.
How opaque the floating panel is. 100% is fully solid; lower values let the desktop blur through.
Profiles
Task-focused personas you switch between. Each profile sets the assistant's name, role, instructions, and how long its replies run. Switch anytime from the panel header.
Manage them in Settings → Profiles.
Start from scratch, duplicate an existing one, or use Create with AI — describe the persona you want (e.g. "a concise release-notes writer" or "a patient tutor who explains every step") and SpeakoFlow drafts it. You can even dictate the description.
- Name & role — shown on the card and in the panel.
- Avatar — upload an image or use the default.
- Instructions — the system prompt that shapes its personality.
- Response length — Short, Medium, Long, or Inherit global to follow your main Assistant setting.
- Greeting — an optional opening line shown in an empty chat.
Share profiles or back them up. Export a profile to a file and import it on another machine. You can restore the built-in profiles at any time.
There's a built-in Cat profile that ignores the AI entirely and just meows. Nothing to configure — it's there for fun.
Personal memory
Optional, on-device memory so the assistant learns how you like to work and remembers useful facts between chats. Off until you turn it on.
Everything here is stored on your device and fully editable in Settings → Memory.
Let the assistant keep and use a personal memory.
For the current conversation, don't use memory and don't learn anything new.
How much memory to add to each reply: Light, Balanced, or Detailed. More detail is richer but a little slower.
What it stores
- About you — a short, always-on summary the assistant keeps in mind (e.g. "Prefers short, direct answers. Works in Rust and TypeScript.").
- Notes — specific things worth remembering, either learned from your chats or added by you. Each note shows a confidence level and whether it was learned or added by you.
Managing memory
- Learn from this chat now — pull durable facts from the current conversation into memory on demand. (Otherwise, learning happens quietly at the end of a conversation.)
- Export / Import — back up or move your memory.
- Wipe memory — delete everything the assistant remembers.
Safety by design. Memory refuses to store secrets, passwords, and personal identifiers, and never lets remembered text override your current message. It's advisory only, stored on-device, and you can view, edit, or erase all of it. Learning runs off the hot path so it never slows a reply.
History
Your recent dictations, Flow generations, and assistant chats — stored on this device, and yours to search, reuse, or delete.
Find it in the History section. Filter by All, Dictations, Flow, or Assistant chats.
What you can do
- Copy a transcription, a Flow output, or a whole conversation.
- Star a dictation to keep it — starred recordings are never auto-deleted.
- Re-transcribe a recording with your current model, e.g. after switching to a more accurate one.
- Continue an assistant chat right where it left off.
- Delete individual entries, or open the recordings folder on disk.
Recording storage
These controls affect dictation and Flow recordings only. Assistant chats have separate storage and are never counted or auto-deleted by these rules.
How unstarred recordings are cleaned up: Never, Keep a set number, or after 3 days, 2 weeks, or 3 months. Changes apply immediately.
The maximum number of unstarred dictation and Flow recordings to keep when using "Keep a set number." Starred recordings are always kept.
General settings
Shortcuts, your microphone, sounds, startup, and how the app looks. Everything in the General section, grouped as you'll find it.
Appearance
A Light, Dark, or System-matched theme for the app and the assistant panel.
Scale text and controls across the app: Small, Default, Large, or Extra large.
Change the language of the SpeakoFlow interface. (This is separate from your speech-recognition language.)
Recording & input
The keyboard shortcuts that start and stop recording and open the assistant. See Keyboard shortcuts for the full list and defaults.
Hold records while the shortcut is pressed; Tap starts with one press and stops with the next.
The key you tap mid-hold to go hands-free. See Hands-free & tap-to-lock.
Speech-recognition language. A specific language can improve accuracy; Auto detects it.
Select your preferred microphone device.
Sounds
Play a sound when recording starts and stops.
Choose the start/stop cue — several themes are bundled (Default, Marimba, Pop, Click, and more), with a preview button.
Which device plays feedback sounds.
Adjust the volume of audio feedback sounds.
Startup & tray
Open SpeakoFlow automatically when you sign in.
Launch to the tray without opening the window.
Keep SpeakoFlow in your system tray. The tray menu offers Home, Check for updates, Copy last transcript, Unload model, and Quit.
Updates
Let SpeakoFlow check for new versions and offer to download them. Turn off to disable automatic checks.
Debug & advanced
Diagnostics and fine-grained knobs. Open the Debug section with Ctrl+Shift+D (Cmd+Shift+D on macOS). Most people never need these.
Where logs are written, and how verbose they are. Raise the level when reporting a bug.
Sensitivity for custom-word corrections — how aggressively SpeakoFlow snaps near-misses to your custom words.
Which microphone to use when the laptop lid is closed.
Mute system audio during recording so playback doesn't bleed into your dictation.
Choose the shortcut backend: handy-keys (supports modifier-only combos like Ctrl+Super) or the Tauri global-shortcut plugin. If a switch makes your shortcuts incompatible, they reset to defaults.
Milliseconds to wait before pasting. Increase if the wrong text is occasionally pasted.
Milliseconds to keep recording after you release the key, so trailing words aren't clipped.
Reveal in-development features (like AI cleanup and live transcription). Off by default.
Command-line flags
SpeakoFlow accepts command-line flags on all platforms, for scripts, window managers, and autostart setups. Remote-control flags are sent to the already-running instance.
| Flag | What it does |
|---|---|
--toggle-transcription | Toggle recording on/off on a running instance. |
--toggle-post-process | Toggle recording with AI cleanup on/off. |
--cancel | Cancel the current operation on a running instance. |
--start-hidden | Launch without showing the main window (tray icon visible). |
--no-tray | Launch without a system tray (closing the window quits the app). |
--debug | Enable debug mode with verbose (Trace) logging. |
CLI flags are runtime-only overrides — they don't modify your saved settings. You can bind them to a hardware key or a window-manager shortcut for a fully custom trigger.
Privacy
Private by default, and honest about it.
- Your voice is transcribed on your device and never uploaded. Speech-to-text always runs locally.
- The assistant only contacts the model provider you choose — which can be a fully local one, with no network at all.
- No telemetry, no account. SpeakoFlow doesn't phone home.
- Optional features stay off until you turn them on — web search and personal memory are opt-in.
- Memory is stored on your device, where you can view, edit, or erase it, and it refuses to keep secrets or personal identifiers.
- API keys live in your OS keychain, not a plain-text file.
Troubleshooting
Fixes for the issues people hit most often.
Microphone access is blocked
- Windows Enable microphone access in Settings → Privacy & security → Microphone, including desktop-app access.
- macOS Grant access in System Settings → Privacy & Security → Microphone.
- Linux Grant access in your system's sound or privacy settings.
Transcribed text won't type macOS
SpeakoFlow needs Accessibility permission to type into apps. Grant it in System Settings → Privacy & Security → Accessibility, then restart the app.
Screen vision does nothing macOS
Grant Screen Recording permission in System Settings → Privacy & Security → Screen Recording and restart SpeakoFlow for it to take effect.
The recording overlay won't stay on top Linux
The overlay has to float above every other window, which on Linux requires either the wlr-layer-shell protocol (wlroots compositors like Sway and Hyprland, and KDE Plasma) or classic X11 "keep above" stacking. Native GNOME/Wayland supports neither.
SpeakoFlow handles this automatically: when it detects GNOME on Wayland it runs under XWayland, where "keep above" works. If you need to override the behavior:
- Force native Wayland anyway (overlay may not stay on top): launch with
SPEAKOFLOW_ALLOW_WAYLAND=1. - If the overlay misbehaves under a layer-shell compositor, disable layer shell with
SPEAKOFLOW_NO_GTK_LAYER_SHELL=1.
The wrong text pastes, or words get clipped
Increase the Paste delay and the Extra recording buffer in Debug settings. If a specific app misbehaves with clipboard paste, try the Direct paste method.
Automatic update failed Windows
Portable installs can't update automatically. Download the latest NSIS installer from GitHub Releases, install it to the same folder, then copy your Data/ folder (settings, models, recordings) from the old version to the new one.
About & license
Where things live, and the people this is built on.
Folders
The About section shows your app-data folder, where SpeakoFlow keeps your models, history, and settings, plus quick links to the model and settings locations.
Tech stack
- App: Tauri 2 with a Rust backend and a React + TypeScript frontend.
- Speech-to-text: whisper.cpp and Parakeet with GPU acceleration, plus Silero VAD for voice detection.
- Assistant: a built-in llama.cpp engine, or any OpenAI-compatible provider you configure.
- Text-to-speech: Kokoro locally, with OpenAI-compatible, ElevenLabs, and Azure options.
License & credits
SpeakoFlow is open source under the MIT License. It started as a fork of Handy by CJ Pais (also MIT), which provides the local dictation core. Thanks also to Tauri, whisper.cpp, llama.cpp, Silero VAD, and Kokoro.
Have an idea or hit a bug not covered here? Open a Discussion or check the repository.