MT server protocol (v1)
July 5, 2026 ยท View on GitHub
alignatt-mt-server serves the AlignAtt MT policy to external ASR frontends
over WebSocket. The client owns transcription; the server owns translation.
This document is the normative protocol spec. WhisperLiveKit
(--translation-backend alignatt) is the reference client.
alignatt-mt-server --preset gemma_low_latency --port 8765
All frames are JSON text. One WebSocket connection = one translation session.
Session lifecycle
- Client connects and sends
init. - Server replies
init_ok(orerror, then closes). - Client streams
updatemessages; the server answers each serviced update with atranslationmessage. - Plain WebSocket close ends the session. The client should send its last
update with
is_final: truebefore closing if it wants an end-of-stream quality pass.
Client to server
init
{
"type": "init",
"protocol_version": 1,
"source_lang": "en",
"target_lang": "de",
"preset": "gemma_low_latency",
"overrides": {"translation_alignatt_inaccessible_ms": 0},
"context_text": "Talk title, glossary, or domain hints (optional)",
"history": [["Previous sentence.", "Vorheriger Satz."]],
"accepted_target_prefix": ""
}
protocol_version(required): must be1.source_lang,target_lang: language codes. The server validates the direction against the calibrated alignment-heads files it ships and returnsunsupported_directionwith thesupportedlist otherwise.preset(optional): server-side runtime preset name.overrides(optional): per-session knobs; only whitelisted keys apply (seeCLIENT_OVERRIDE_WHITELISTinmt_server.py), unknown keys are ignored silently.context_text(optional): free-text domain context injected into the MT prompt outside the source-marked region (it is never translated).history(optional):[source, target]pairs restored after a reconnect so translation-history windows keep working.accepted_target_prefix(optional): the partial target text already shown to the user for the open utterance before a reconnect. The server continues from it and never re-emits or contradicts it.
update
{
"type": "update",
"seq": 42,
"utterance_id": 3,
"words": [["Hello", 120.0, 400.0], ["world,", 520.0, 900.0]],
"tail": {"words": [["how", null, null], ["are", null, null]]},
"clock_ms": 2350.0,
"is_final": false
}
- Updates carry the FULL state of the open utterance, not deltas:
wordsis every committed word of the utterance, in order, as[text, start_ms, end_ms](timestamps may benull). A newer update supersedes an older unserviced one losslessly, which is what makes server-side coalescing and reconnects trivial. tail.words(optional): the ASR's unstable hypothesis tail. The MT model drafts over it, but AlignAtt only commits target tokens whose attention lands on committed words. Feed it when the upstream ASR is append-only (e.g. the qwen3 causal backend); omit it for hypothesis-churning backends.clock_ms: monotone stream clock (the audio head position). Must be at least the largest committedend_ms.is_final: truecloses the utterance: the server runs a full-quality translation of the whole utterance (reusing the streamed partial as a prefill) and the next update must useutterance_id + 1.seq: client-chosen monotone integer, echoed back. Because partial updates coalesce (latest wins), someseqvalues never get a response.is_finalupdates are sticky and always processed, in order.
Server to client
init_ok
{"type": "init_ok", "protocol_version": 1, "direction": "en-de",
"preset": "gemma_low_latency", "model": "gemma_vllm_alignatt"}
translation
{
"type": "translation",
"seq": 42,
"utterance_id": 3,
"final": false,
"committed_text": "Hallo Welt,",
"committed_delta": " Welt,",
"buffer_text": "wie",
"covered_source_units": 2,
"stop_reason": "alignatt:commit_fast_path"
}
committed_text: cumulative accepted target for the open utterance. Append-only across the utterance: each value extends the previous one (previous + committed_delta == committed_text, delta verbatim including leading whitespace).buffer_text: the model's draft beyond the accepted prefix (display-only, may be rewritten at any time).covered_source_units: 1 + the highest source word index the accepted target attends to (from the calibrated alignment heads);nullwhen unavailable. Clients use it to timestamp translation segments.- When
final: true,committed_textis the full-quality translation of the closed utterance and replaces the streamed partial at the line level;forced_final: trueadditionally signals a server-side rollover (--max-utterance-wordsexceeded), after which the client must bump itsutterance_idexactly as for its own finalization.
error
{"type": "error", "code": "unsupported_direction",
"message": "...", "supported": ["en-de", "en-it", "en-zh"]}
Codes: bad_protocol, unsupported_direction, busy (server at
--max-sessions), processing_failed (that update was dropped; the session
stays usable).
Scheduling and backpressure
The server serializes all MT work on one engine (the vLLM attention observer
supports a single in-flight request). Per connection it keeps a mailbox of
size one for partial updates (latest wins) plus a sticky ordered queue for
finals. Clients should self-pace: coalesce commits, avoid calling per single
word, and treat translation responses as the natural pacing signal.
Frontier semantics (why this protocol is shaped like this)
build_source_accessibility_frontier marks committed words accessible
(they carry real timestamps) and tail words inaccessible (they carry none).
The prompt text itself does not encode accessibility, so when the ASR
commits words that were already in the prompt as tail, the prompt bytes do
not change; with translation_alignatt_commit_fast_path (default on for
this server) the previously held target tokens are released by re-cutting
the cached draft, without a new MT engine call. Feeding the tail therefore
converts the ASR commit latency into MT speculation time.