Personalization and Memory

September 5, 2026 · View on GitHub

The voice frontend composes four context layers. The first two define the assistant; the last two describe the current user and can be supplied by a replaceable memory provider.

LayerSourceResponsibility
Core policyconfig/frontend-agent/PROMPT.mdTool protocol, permission, safety, and task boundaries; user memory cannot override it
Assistant profileASSISTANT.mdInstance-wide default identity, personality, relationship stance, and expression style; configured by users or downstream products
User preferencesuser (default provider: USER.md)Explicit long-term personalization for the current user; overrides the default persona
Long-term memorymemory (default provider: MEMORY.md)Durable facts and decisions used to understand the user and answer questions; no behavioral authority

Instruction conflicts resolve in this order: core policy, the user's current explicit request, the user preferences, then the assistant profile. Long-term memory is not part of the instruction hierarchy; it is evidence only, and the user's current statement wins when facts conflict. Saying “keep replies shorter from now on” or “call yourself Skiff from now on” updates the current user's USER.md, not instance-wide ASSISTANT.md; a temporary request applies only to the current turn.

Default implementation

With the default configuration, the Gateway uses its built-in Markdown provider. It keeps the following files under the configuration directory (~/.config/qwaudio/ for the CLI):

FileDescription
ASSISTANT.mdInstance-wide default persona: identity, personality, relationship stance, and expression style
USER.mdLong-term personalization overlay for the current user
MEMORY.mdDurable facts and decisions about the user
memory-audit.jsonlDiagnostic log for automatic memory patches, skips, and failures

These files remain local and are never committed to the source repository. USER.md and MEMORY.md are the default provider's physical representation, not a requirement imposed on other providers.

Assistant Profile

On first launch, the packaged config/frontend-agent/ASSISTANT.md template is copied to the local ASSISTANT.md; upgrades never overwrite it. Edit the local file to change the whole assistant instance's default name, personality, relationship stance, and expression style. Changes apply to the next voice session. You can also point QWEN_AUDIO_AGENT_ASSISTANT_PROFILE_PATH to another file.

ASSISTANT.md is neither conversation memory nor runtime policy. The assistant never changes it through the memory tool. Statements about tools, permissions, safety, memory, task routing, or capabilities cannot override PROMPT.md.

User Preferences

USER.md is the current user's long-term personalization overlay on the default persona, not a second assistant persona or a general fact store. It may contain how the assistant addresses the user, how this user addresses the assistant, and explicitly requested language, reply style, and default behavior. It changes only for an explicit user setting or correction. Session-end reconciliation may recover such an explicit directive, but it never infers one.

Classify by scope, not grammatical subject. “The assistant's default name is Qwen Audio” belongs in ASSISTANT.md; when the current user says “call yourself Skiff from now on,” Skiff is that user's override and belongs in USER.md. Likewise, “continue project A by default” belongs in USER.md, while “project A uses React” is a fact for MEMORY.md. It is ordinary Markdown. Tool writes take effect immediately; direct edits apply to the next voice session. To store it elsewhere, set QWEN_AUDIO_AGENT_USER_MODEL_PATH (the legacy QWEN_AUDIO_AGENT_USER_PROFILE_PATH name is still accepted).

Do not store passwords, API Keys, verification codes, or tokens in this file.

Legacy profile, rules, and user records from frontend-memory.json are migrated into USER.md on first launch.

Preference self-update (default provider only, off by default)

With QWEN_AUDIO_PREFERENCE_LEARNING=on, the Gateway observes user traits from a finished session and writes them to USER.md only after enough cross-session confirmation. It is off by default because it adds one model call per session.

Four fields only, with deliberately narrow value spaces:

FieldMeaning
occupationOccupation
special_skillsTechnologies or domains the user is strong in, capped at 6
response_lengthReply length; only brief, normal, or detailed
response_styleReply style

Writes land in the ## 观察推断 (observed) section of USER.md, kept physically separate from ## 用户明确要求 (explicitly stated). An explicit statement always wins on conflict. The split prevents inferred content from polluting what the user wrote: the user can see which lines came from their own words and which the system guessed, and can edit or delete the latter.

Promotion gate

An observation reaches the document only when confirm ≥ 2 and the confirmations come from ≥ 2 distinct sessions. confirm resets after 90 days with no new confirmation.

Four structural guards

A model sometimes supplies a genuine quote while the inference from that quote does not hold. Repeated sampling cannot filter this class out — the user says the same sentence every session, the model makes the same wrong inference, and the counter climbs to the threshold anyway. The guards therefore apply at admission time:

GuardWhat it blocks
quote_not_from_userThe quote must appear verbatim in a user turn: blocks fabricated evidence, assistant speech treated as user preference, and self-reinforcement
value_not_anchoredThe literal parts of the conclusion must be findable in the quote
value_parrots_quoteValue equals the quote — that is repetition, not feature extraction
quote_not_about_interactionFor interaction-preference fields the quote must address the assistant: blocks "how long the content should be" being read as "how long the reply should be"

Diagnostics land in memory-audit.jsonl, so a rejected observation can be explained after the fact.

Replacing the memory implementation

A host application can inject a versioned MemoryProvider to replace the user and memory layers. The provider may use Markdown, a database, a remote service, or a semantic memory engine. It owns storage and retrieval; when it declares sessionObservation, it also owns session-end learning and the built-in extractor and preference learner are disabled. Providers that need acoustic context may additionally opt into bounded PCM16 observation; transcript-only providers and the default Markdown implementation receive no microphone audio.

This boundary deliberately excludes PROMPT.md and ASSISTANT.md. Core behavior and the instance-wide default persona therefore stay stable regardless of the selected memory system. The Realtime Agent continues to use the same memory tool and the same logical user / memory scopes.

See Long-Term Memory for the provider contract and the VoiceMem setup example for the optional built-in integration.

  • Long-Term Memory — the default Markdown implementation, the memory tool, session digests, and the replaceable provider contract