Skip to main content
Tune how the assistant carries a conversation, in Settings → Assistant → Behavior.

Conversation memory

How many recent messages are sent as context each turn. Higher keeps more of the thread in mind at the cost of a bigger prompt. 0 disables it and treats every message as standalone.

Summarize long chats

On by default. Once the un-summarized part of a conversation grows past your Conversation memory window, the oldest of those messages are folded into a rolling summary instead of being dropped. The six most recent messages stay verbatim. The summary is injected as a context note ahead of the history, so a long chat keeps its beginning in condensed form rather than forgetting it. Turn it off and a long conversation simply loses its oldest messages once they fall outside the window.

Response length

No effect by default, which adds nothing and leaves your prompt exactly as written. The other three append a length instruction to the system prompt at request time. Each option tells the model to still match your intent, so a greeting gets a short friendly reply even on Long rather than a padded one. Individual profiles can override this, and the profile’s choice wins for its own turns.
Cloud providers get exactly the number of messages you ask for. Azure and local engines (the built-in one, or an Ollama or LM Studio server on loopback) also get a character cap on top.Azure’s gateway rejects oversized request bodies, and a local model runs a small context window, so sending your full history there would fail rather than work slowly. Turns carrying a screenshot trim harder on those same providers, since the image already dominates the budget.
The trigger is the message-history window above, not the model’s context window.The fold runs in the background after a reply finishes, so it never delays an answer, and a subtle indicator shows while it is working. If it fails, the summary is left alone and the next turn tries again.
You can compact a conversation on demand by typing /summarize in the panel. That one replaces the visible transcript with its summary, rather than working quietly in the background.
This is short-term, per-conversation context. For durable facts the assistant remembers across chats, see Personal memory, a separate opt-in feature.
Last modified on August 7, 2026