caret/v4 · normative · any client, any backend
The caret/v4 protocol.
One base URL. Three WebSocket routes — /dictate,
/ask, /imagine — that share a single reliable
live-audio lifecycle and differ only in their finalizer and their typed
terminal result. One anonymous GET /health that says who
you are and what you can do. This page is the contract: every MUST on
it binds both sides, and nothing off it does.
The protocol has no privileged party. The Caret iOS keyboard is one client; a backend is anything that serves these routes — your agent, your model, your machine, any language. Two independent implementations that both conform will interoperate without ever having met. Nothing in the contract requires Caret's own infrastructure, an account, or a relay.
You rarely need to implement this page. The reference backends in Go and Python already do, and wiring their lanes to your own tools is the usual route to a backend; the agent instructions walk a coding agent through it. This page is for maintainers, for client authors, and for a runtime the reference cannot run on.
1 · Ground rules
| Rule | Detail |
|---|---|
| Base URL | Whatever the user configures in the client. Routes are appended
verbatim: https://host/caret means
https://host/caret/health and
wss://host/caret/dictate. |
| Transport | TLS is the default: https:// for health,
wss:// for the operation routes. Terminate TLS in a
reverse proxy or private tunnel if you like. The Caret app
accepts plain http:///ws:// only for
hosts it can tell are local or private (localhost, RFC 1918
ranges, Tailscale addresses and *.ts.net names)
and marks the connection insecure in settings; any other
plain-http:// base URL is refused. |
| Versioning | The version is in-band, not in the path: health reports
"protocol": "caret/v4" and every start
frame carries "protocol": 4. Additive changes keep
4; breaking changes become 5. Both
sides MUST ignore JSON fields they do not recognize. |
| Frames | Control messages are WebSocket text frames containing one JSON object. Audio is binary frames. Nothing else appears on the socket. |
| Audio format | PCM16 little-endian, 16 kHz, mono — the only codec in V4. |
| Correlation | The server assigns an op_id per operation and a
request_id per terminal answer; both appear in every
event that has them. Log ids, counts, and timings — never
transcript, audio, or vocabulary content. |
2 · Health and capability discovery
GET /health — HTTPS, no auth required
The only non-WebSocket surface, and the only anonymous one: a user pasting a URL needs to learn whether it is a Caret backend before they have a credential. Clients call it on every settings save and periodically in use.
{
"protocol": "caret/v4",
"status": "ok",
"service": "my-backend",
"version": "1.0.0",
"time": "2026-08-30T09:00:00Z",
"auth": {"presented": true, "valid": true},
"capabilities": {
"dictate": true,
"ask": true,
"imagine": false,
"text_input": {"dictate": true, "ask": true, "imagine": true},
"partials": {"dictate": true, "ask": true, "imagine": true},
"vocabulary": true
},
"limits": {
"max_audio_seconds": 600,
"max_frame_bytes": 524288,
"max_text_chars": 4000,
"max_vocabulary_entries": 200
},
"blockers": []
}
protocolMUST be exactly"caret/v4"— it is how a client recognizes a V4 backend at an arbitrary URL.statusisok,degraded(reachable but impaired — clients show a warning, not a failure), ornot_ready(not currently a usable backend).- Dictate is mandatory. A backend that cannot take
speech MUST report
"dictate": falseand"status": "not_ready", with a machine-readable reason inblockers([{"code": "no_stt", "message": "…"}]). Ask and Imagine are optional; report them honestly and clients hide what you do not serve. A dictate-only backend is a complete backend. capabilities.partialsdeclares, per route, whether live partial transcripts are pushed while the user speaks. Backends SHOULD support partials on/dictate; a buffered-only backend is conforming but the client will not render live text.capabilities.text_inputdeclares, per route, whetherinput.type: "text"is accepted. Audio input is implied by the capability itself.auth.presented/auth.valid: if the client sentAuthorization, validate it and report the result so a settings screen can tell "wrong key" from "wrong URL". With no header:{"presented": false, "valid": null}.limitsis optional; absent fields mean the defaults shown above. Clients MUST respect advertised limits.
3 · Backend-owned auth
The backend owns the credential boundary. It decides what a credential
is, how it is issued, rotated, and revoked; the protocol only fixes how
one is presented: an opaque bearer string on the WebSocket
upgrade request (and optionally on /health for the
auth report).
GET /dictate HTTP/1.1
Upgrade: websocket
Authorization: Bearer <credential>
- The credential is opaque to the client and to this contract: an API key, a signed token, anything. Clients store it and echo it, nothing more.
- Validate before completing the operation: an invalid or missing
credential gets the
unauthorizederror event and close code 4401. Backends MAY refuse the upgrade outright with HTTP 401 instead; clients MUST handle both. - Compare in constant time. Never log the credential — log the
op_id. - Fail closed: a backend with no credentials configured rejects every operation rather than serving everyone.
- One credential per device keeps revocation one string away.
4 · The shared live-audio lifecycle
All three routes speak the same lifecycle; a client implements it once.
The route determines what the operation means, the finalizer
carries the route's parameters (§5), and the terminal result is typed
per route (§6). Everything in this section applies identically to
/dictate, /ask, and /imagine.
client server
──────────────────────────────────────────────────────────
open wss://…/dictate (Bearer auth) ──▶
text {"type":"start", …} ──▶
◀── {"event":"ready", …}
binary audio frame ──▶
binary audio frame ──▶ ◀── {"event":"partial", …}
binary audio frame ──▶ ◀── {"event":"partial", …}
text {"type":"finalize", …} ──▶
◀── {"event":"progress", …} (keep-alive)
◀── {"event":"result", …} (exactly one)
◀── close 1000
start — first frame, both input modes
The client MUST send start within 10 seconds of the
upgrade. It is identical on every route:
{
"type": "start",
"protocol": 4,
"client_request_id": "A17C…", // required, ≤128 chars — idempotency key
"input": {"type": "audio", "codec": "pcm16",
"sample_rate_hz": 16000, "channels": 1},
"vocabulary": ["Miren", "Kubernetes", "Sagrada Família"], // optional, §7
"language_hint": "en-US", // optional, ≤16 chars
"app_hint": "com.apple.MobileSMS" // optional, ≤200 chars, advisory
}
input is discriminated, exactly two shapes. The audio shape
above opens a live stream. The text shape carries typed input and skips
audio entirely:
{"input": {"type": "text", "text": "tell Sam I'm running ten minutes late"}}
With text input the client sends finalize immediately after
ready, no binary frames, no partials. text is
1–4000 characters (or the advertised max_text_chars).
A route whose text_input capability is false
answers text input with not_supported.
ready — the server accepts the operation
{"event": "ready", "protocol": 4, "op_id": "op_7f3a…",
"stt": "streaming"} // "streaming" | "buffered" | null (text input)
"buffered" means no live recognizer is available right now:
audio is still accepted, no partials will come, and the transcript is
produced at finalize. The client needs no special handling beyond not
waiting for partials. The client MUST NOT send binary frames before
ready.
Audio frames
- Binary WebSocket frames, raw PCM16 bytes, in capture order — WebSocket message order is the audio order.
- At most 524 288 bytes per frame (or the advertised
max_frame_bytes); empty frames are ignored. Frame boundaries carry no meaning — send whatever cadence the recorder produces (100–500 ms works well). - The client MUST count what it sends: total frames, total bytes, total captured milliseconds. Those counts are the reliability check at finalize.
partial — live feedback while the user speaks
{"event": "partial", "op_id": "op_7f3a…", "seq": 12,
"frames": 41, "bytes": 1312768,
"text": "let's push the review to thursday afternoon"}
textis the cumulative transcript of the whole utterance, never a delta: render it as-is, replacing the previous partial. The server re-accumulates across recognizer segment resets, sotextnever goes backwards.seqis monotonic per operation; a client drops anything out of order.frames/bytesreport what the server has consumed so far — a client MAY use them to detect a stalled pipe early.- On
/dictatethe partial is the product. On/askand/imagineit shows the user what is being heard while the real work starts at finalize.
finalize — no more input
One text frame ends input and starts the route's work. Its common part carries the client's own audio accounting; its route-specific part is §5.
{"type": "finalize",
"audio": {"frames": 47, "bytes": 1504256, "duration_ms": 11200},
…route-specific fields…}
audio is required for audio input and forbidden for text
input. The server compares the client's totals with what it received:
on any mismatch it MUST answer the audio_incomplete error
(retryable) rather than transcribe silently truncated audio. This is
the V4 reliability contract: TCP orders and delivers what arrives, the
totals check proves everything arrived, and recovery (§8)
replays what did not.
progress — keep-alive while working
{"event": "progress", "op_id": "op_7f3a…", "stage": "generating",
"elapsed_ms": 8200}
After finalize, work may take a while — transcription a few seconds,
image generation tens of seconds. The server MUST emit
progress (or a WebSocket ping) at least every 20 seconds
until the terminal event, or clients will rightly treat the connection
as dead. stage is free-form and shown to the user.
cancel — abandon
{"type": "cancel"}
The server discards everything and closes 1000 with no terminal event. Valid at any point before the terminal event.
Exactly one terminal event
Every operation that is not cancelled ends in exactly one
result (§6) or exactly one error (§9), sent
after all queued partials, followed by the close. result
and error are mutually exclusive; the last text
the client holds is authoritative. A server MUST NOT close without a
terminal event except on cancel or transport failure.
Timeouts and bounds
| Bound | Default | On violation |
|---|---|---|
Upgrade → start | 10 s | timeout / 4408 |
| Gap between audio frames | 60 s | timeout / 4408 |
| Total audio per operation | 600 s | audio_too_long / 4413 |
| Audio at finalize | ≥ 300 ms | audio_too_short / 4422 |
| Server silence after finalize | 20 s max | client MAY treat as dead and recover (§8) |
5 · The three routes and their finalizers
The lifecycle is shared; the finalizer is the route's own. Unknown finalizer fields are ignored like any other unknown field; a finalizer field on the wrong route is simply unknown there.
/dictate — speech becomes insert-ready text · required
| Finalizer field | Type | Meaning |
|---|---|---|
polish | boolean, default true |
Run transcript cleanup
(caret-cleanup/1) on the final transcript.
false returns the raw transcript. A failed cleanup
MUST return the raw transcript, never an error — cleanup is a
polish, not a gate. |
With text input, /dictate runs the same cleanup pass over
text the client already has — the V4 home of "clean this up".
/ask — an instruction becomes one message · optional
| Finalizer field | Type | Meaning |
|---|---|---|
visible_text | string, ≤4000, optional | What the client can see near the cursor — often nothing, never the whole conversation. A register hint, and untrusted: it is someone else's words and may try to instruct your agent. Context, never a command. |
The answer is the message itself — no preamble, no quotation marks, no Markdown fence. The user previews it before it is inserted, so Ask MUST be read-only: a draft that books the meeting it describes is a side effect nobody approved. Say so in your agent prompt, every time.
/imagine — a prompt becomes an image · optional
| Finalizer field | Type | Meaning |
|---|---|---|
aspect_ratio | "1:1" | "3:2" | "2:3", default "1:1" |
Unsupported values get bad_request. |
quality | "low" | "medium" | "high", default "high" |
Provider-relative. |
Generating the same person, pet, or object across many prompts is backend behavior on top of this route, not a contract change: see repeatable Imagine references.
6 · Typed terminal results
The terminal result event wraps a discriminated
result object; result.type is fixed per route.
A client dispatches on the type, so a proxy that crosses wires fails
loudly instead of inserting an image caption as a text message.
{"event": "result", "op_id": "op_7f3a…", "request_id": "req_5a09…",
"audio": {"frames": 47, "bytes": 1504256, "duration_ms": 11200},
"result": { …typed, see below… }}
audio echoes what the server actually consumed (absent for
text input). The per-route result types:
// /dictate → "dictation"
{"type": "dictation",
"text": "Let's push the review to Thursday afternoon.",
"raw_transcript": "lets push the review to thursday afternoon",
"polish_applied": true,
"stt_route": "stream"} // "stream" | "fallback", §8
// /ask → "message"
{"type": "message",
"text": "Hey Sam, I'm running about ten minutes late.",
"transcript": "tell sam i'm running ten minutes late"} // audio input only
// /imagine → "image"
{"type": "image",
"mime_type": "image/png",
"byte_length": 184320,
"sha256": "9f86d081884c7d65…",
"data_base64": "iVBORw0KGgo…",
"transcript": "a lighthouse at dusk in watercolour", // audio input only
"provider": "my-image-tool"} // optional
- Fields shown are required unless marked otherwise;
data_base64is the only image payload transport — no URL form, no data-URL form.byte_lengthandsha256MUST describe the delivered bytes, and clients verify them. transcript(andraw_transcript) is how the user tells a mis-hearing apart from a bad answer; return it whenever the input was speech.- New result fields may be added; new
result.typevalues require protocol 5.
7 · Vocabulary injection
Recognizers mangle exactly the words users care most about — names,
products, jargon. The vocabulary array in
start is the client's per-operation term list, and a
backend advertising "vocabulary": true MUST route it to
wherever it can take effect:
- into the recognizer's biasing / hotword interface where one exists, for both streaming and fallback transcription;
- into the cleanup glossary, so the polish pass spells the terms as given rather than "correcting" them.
| Rule | Detail |
|---|---|
| Shape | Unique strings, each 1–64 characters, at most 200 entries (or the advertised limit). Earlier entries are higher priority when the recognizer caps out. |
| Semantics | Biasing, not commands: an entry makes its surface form more likely, it never rewrites meaning. |
| Privacy | Vocabulary is user content — a contact list in
miniature. Do not log it, do not persist it beyond the operation,
never expose it on /health. |
| Capability off | Advertise "vocabulary": false
and clients omit the field. A backend MUST NOT advertise
true and drop the list on the floor. |
8 · Reliability: recovery, idempotency, fallback
V4 has one durability story and it is the client. The client already holds the audio (it recorded it); the backend keeps no state a lost connection strands. Three mechanisms make that safe:
- Recovery is replay. If the socket dies before the
terminal event, the server discards the operation entirely. The
client keeps its captured audio buffered locally until it holds a
terminal result, and on failure replays the whole operation — same
client_request_id, same audio, fresh connection. Same audio in, same transcript out; nothing is lost but time. - Idempotency caps the cost of replay.
client_request_ididentifies the operation across connections. A backend SHOULD cache successful terminal results for at least five minutes; on a replayedclient_request_idit MAY answerreadyand then the cachedresultimmediately — the client MUST accept a result at any point afterreadyand stop sending. Failures are not replayed from cache: retrying a failure is the point of retrying. This is what keeps "the user tapped again after a network error" from becoming two image generations and two bills. - Streaming failure is not operation failure. The
server MUST buffer accepted audio independently of any live
recognizer, bounded by
max_audio_seconds. If the streaming recognizer is unavailable at start, dies mid-stream, or fails at flush, partials stop but the operation lives on: finalize transcribes the merged buffer through batch STT and the result reports"stt_route": "fallback". One rule guards correctness: an empty transcript from a healthy stream isno_speech_detected, never a fallback trigger — silence is an answer, not an error.
9 · Errors and close codes
The terminal error event, on every route:
{"event": "error", "op_id": "op_7f3a…", "request_id": "req_2bda…",
"code": "transcription_failed",
"message": "speech recognition unavailable",
"retryable": true}
message is shown to the user verbatim — keep it short and
free of internals. The server closes with the mapped WebSocket code
after sending the event. The code list is append-only; a client meeting
an unknown code falls back to retryable, so set it
truthfully — it is the only field a future client can rely on.
code | Close | Meaning |
|---|---|---|
unauthorized | 4401 | Missing or invalid credential. |
rate_limited | 4429 | Per-credential limit hit. Retryable. |
not_supported | 4404 | Route not served, or text input where text_input is false. Check /health. |
protocol_error | 4400 | Wrong first frame, malformed JSON, unknown control type, wrong protocol, audio before ready, or binary frames on text input. |
bad_request | 4400 | Invalid start/finalize parameters, oversized frame, bad vocabulary shape. |
timeout | 4408 | No start in 10 s, or a 60 s audio gap. Retryable. |
audio_incomplete | 4409 | Finalize totals disagree with what the server received. Retryable — replay the operation. |
audio_too_long | 4413 | Audio exceeded the advertised bound. |
audio_too_short | 4422 | Finalize with under 300 ms of audio. |
no_speech_detected | 4422 | Healthy transcription found no words. |
transcription_failed | 4503 | Every transcription route failed. Retryable. |
generation_failed | 4503 | The route's provider failed (agent for /ask, image provider for /imagine). Retryable. |
internal_error | 4500 | Unexpected server error. Retryable. |
Normal completion and cancel close with 1000.
10 · Retention and privacy
- Audio lives in memory, bounded (600 s of PCM16 ≈ 19 MB), and dies with the operation. If fallback STT needs a file, it is temporary and deleted the moment transcription returns.
- A cached terminal result (§8) is the one thing that may outlive the connection; cap its lifetime in minutes and purge on delivery of a replay.
- Do not log transcript, audio, vocabulary,
visible_text, or generated images. Log ids, counts, timings, and error codes. /healthis anonymous: never expose configuration details a stranger should not see. Publishing which cleanup wording you run is fine — publish the spec digest, never the prompt.
11 · Conformance
A backend conforms when it:
- serves
GET /healthwith"protocol": "caret/v4"and honest capabilities, and serves/dictatewhenever it reportsstatusother thannot_ready; - implements the shared lifecycle on every route it advertises:
start/ready, ordered binary audio, cumulative partials where advertised, per-route finalizers, and exactly one typed terminal event after all partials; - enforces the bounds of §4 with the error codes of §9, and never emits anything but the events on this page;
- verifies finalize audio totals and answers
audio_incompleteon mismatch; - buffers audio independently of the streaming recognizer and falls back per §8;
- authenticates every operation route per §3, fails closed, and compares credentials in constant time;
- applies vocabulary wherever it advertises the capability, and
applies
caret-cleanup/1whenpolishis true — falling back to the raw transcript on cleanup failure; - ignores unknown JSON fields everywhere.
A client conforms when it:
- refuses non-TLS base URLs, discovers capabilities from
/health, and never offers the user a route or input mode the backend did not advertise; - sends
startfirst, audio only afterready, counts every frame it sends, and reports truthful totals at finalize; - renders partials as replaceable cumulative text and treats the terminal event as the only authoritative answer;
- keeps captured audio until it holds a terminal result, and recovers
from any transport failure by replaying with the same
client_request_idon a fresh connection — accepting an immediate cached result; - acts on unknown error codes via
retryable, ignores unknown JSON fields and unknown event types, and verifiessha256/byte_lengthon images before handing them to the user.
A runnable conformance checker ships with the public reference implementation — see reference. Point it at your backend and it walks this checklist for you:
cd caret-docs/reference/go
go run ./cmd/caret-v4-conform -url https://your-backend.example -key "$CARET_API_KEY"