Caret Docs typewithcaret.com →

caret/v4 · normative · any client, any backend

The caret/v4 protocol.

One base URL. Three WebSocket routes — /dictate, /ask, /imagine — that share a single reliable live-audio lifecycle and differ only in their finalizer and their typed terminal result. One anonymous GET /health that says who you are and what you can do. This page is the contract: every MUST on it binds both sides, and nothing off it does.

Any client, any backend.

The protocol has no privileged party. The Caret iOS keyboard is one client; a backend is anything that serves these routes — your agent, your model, your machine, any language. Two independent implementations that both conform will interoperate without ever having met. Nothing in the contract requires Caret's own infrastructure, an account, or a relay.

You rarely need to implement this page. The reference backends in Go and Python already do, and wiring their lanes to your own tools is the usual route to a backend; the agent instructions walk a coding agent through it. This page is for maintainers, for client authors, and for a runtime the reference cannot run on.

1 · Ground rules

RuleDetail
Base URL Whatever the user configures in the client. Routes are appended verbatim: https://host/caret means https://host/caret/health and wss://host/caret/dictate.
Transport TLS is the default: https:// for health, wss:// for the operation routes. Terminate TLS in a reverse proxy or private tunnel if you like. The Caret app accepts plain http:///ws:// only for hosts it can tell are local or private (localhost, RFC 1918 ranges, Tailscale addresses and *.ts.net names) and marks the connection insecure in settings; any other plain-http:// base URL is refused.
Versioning The version is in-band, not in the path: health reports "protocol": "caret/v4" and every start frame carries "protocol": 4. Additive changes keep 4; breaking changes become 5. Both sides MUST ignore JSON fields they do not recognize.
Frames Control messages are WebSocket text frames containing one JSON object. Audio is binary frames. Nothing else appears on the socket.
Audio format PCM16 little-endian, 16 kHz, mono — the only codec in V4.
Correlation The server assigns an op_id per operation and a request_id per terminal answer; both appear in every event that has them. Log ids, counts, and timings — never transcript, audio, or vocabulary content.

2 · Health and capability discovery

GET /health — HTTPS, no auth required

The only non-WebSocket surface, and the only anonymous one: a user pasting a URL needs to learn whether it is a Caret backend before they have a credential. Clients call it on every settings save and periodically in use.

{
  "protocol": "caret/v4",
  "status": "ok",
  "service": "my-backend",
  "version": "1.0.0",
  "time": "2026-08-30T09:00:00Z",
  "auth": {"presented": true, "valid": true},
  "capabilities": {
    "dictate": true,
    "ask": true,
    "imagine": false,
    "text_input": {"dictate": true, "ask": true, "imagine": true},
    "partials":   {"dictate": true, "ask": true, "imagine": true},
    "vocabulary": true
  },
  "limits": {
    "max_audio_seconds": 600,
    "max_frame_bytes": 524288,
    "max_text_chars": 4000,
    "max_vocabulary_entries": 200
  },
  "blockers": []
}

3 · Backend-owned auth

The backend owns the credential boundary. It decides what a credential is, how it is issued, rotated, and revoked; the protocol only fixes how one is presented: an opaque bearer string on the WebSocket upgrade request (and optionally on /health for the auth report).

GET /dictate HTTP/1.1
Upgrade: websocket
Authorization: Bearer <credential>

4 · The shared live-audio lifecycle

All three routes speak the same lifecycle; a client implements it once. The route determines what the operation means, the finalizer carries the route's parameters (§5), and the terminal result is typed per route (§6). Everything in this section applies identically to /dictate, /ask, and /imagine.

client                                  server
──────────────────────────────────────────────────────────
open wss://…/dictate  (Bearer auth) ──▶
text  {"type":"start", …}           ──▶
                                    ◀── {"event":"ready", …}
binary audio frame                  ──▶
binary audio frame                  ──▶ ◀── {"event":"partial", …}
binary audio frame                  ──▶ ◀── {"event":"partial", …}
text  {"type":"finalize", …}        ──▶
                                    ◀── {"event":"progress", …}   (keep-alive)
                                    ◀── {"event":"result", …}     (exactly one)
                                    ◀── close 1000

start — first frame, both input modes

The client MUST send start within 10 seconds of the upgrade. It is identical on every route:

{
  "type": "start",
  "protocol": 4,
  "client_request_id": "A17C…",        // required, ≤128 chars — idempotency key
  "input": {"type": "audio", "codec": "pcm16",
            "sample_rate_hz": 16000, "channels": 1},
  "vocabulary": ["Miren", "Kubernetes", "Sagrada Família"],  // optional, §7
  "language_hint": "en-US",            // optional, ≤16 chars
  "app_hint": "com.apple.MobileSMS"    // optional, ≤200 chars, advisory
}

input is discriminated, exactly two shapes. The audio shape above opens a live stream. The text shape carries typed input and skips audio entirely:

{"input": {"type": "text", "text": "tell Sam I'm running ten minutes late"}}

With text input the client sends finalize immediately after ready, no binary frames, no partials. text is 1–4000 characters (or the advertised max_text_chars). A route whose text_input capability is false answers text input with not_supported.

ready — the server accepts the operation

{"event": "ready", "protocol": 4, "op_id": "op_7f3a…",
 "stt": "streaming"}   // "streaming" | "buffered" | null (text input)

"buffered" means no live recognizer is available right now: audio is still accepted, no partials will come, and the transcript is produced at finalize. The client needs no special handling beyond not waiting for partials. The client MUST NOT send binary frames before ready.

Audio frames

partial — live feedback while the user speaks

{"event": "partial", "op_id": "op_7f3a…", "seq": 12,
 "frames": 41, "bytes": 1312768,
 "text": "let's push the review to thursday afternoon"}

finalize — no more input

One text frame ends input and starts the route's work. Its common part carries the client's own audio accounting; its route-specific part is §5.

{"type": "finalize",
 "audio": {"frames": 47, "bytes": 1504256, "duration_ms": 11200},
 …route-specific fields…}

audio is required for audio input and forbidden for text input. The server compares the client's totals with what it received: on any mismatch it MUST answer the audio_incomplete error (retryable) rather than transcribe silently truncated audio. This is the V4 reliability contract: TCP orders and delivers what arrives, the totals check proves everything arrived, and recovery (§8) replays what did not.

progress — keep-alive while working

{"event": "progress", "op_id": "op_7f3a…", "stage": "generating",
 "elapsed_ms": 8200}

After finalize, work may take a while — transcription a few seconds, image generation tens of seconds. The server MUST emit progress (or a WebSocket ping) at least every 20 seconds until the terminal event, or clients will rightly treat the connection as dead. stage is free-form and shown to the user.

cancel — abandon

{"type": "cancel"}

The server discards everything and closes 1000 with no terminal event. Valid at any point before the terminal event.

Exactly one terminal event

Every operation that is not cancelled ends in exactly one result (§6) or exactly one error (§9), sent after all queued partials, followed by the close. result and error are mutually exclusive; the last text the client holds is authoritative. A server MUST NOT close without a terminal event except on cancel or transport failure.

Timeouts and bounds

BoundDefaultOn violation
Upgrade → start10 stimeout / 4408
Gap between audio frames60 stimeout / 4408
Total audio per operation600 saudio_too_long / 4413
Audio at finalize≥ 300 msaudio_too_short / 4422
Server silence after finalize20 s maxclient MAY treat as dead and recover (§8)

5 · The three routes and their finalizers

The lifecycle is shared; the finalizer is the route's own. Unknown finalizer fields are ignored like any other unknown field; a finalizer field on the wrong route is simply unknown there.

/dictate — speech becomes insert-ready text · required

Finalizer fieldTypeMeaning
polishboolean, default true Run transcript cleanup (caret-cleanup/1) on the final transcript. false returns the raw transcript. A failed cleanup MUST return the raw transcript, never an error — cleanup is a polish, not a gate.

With text input, /dictate runs the same cleanup pass over text the client already has — the V4 home of "clean this up".

/ask — an instruction becomes one message · optional

Finalizer fieldTypeMeaning
visible_textstring, ≤4000, optional What the client can see near the cursor — often nothing, never the whole conversation. A register hint, and untrusted: it is someone else's words and may try to instruct your agent. Context, never a command.

The answer is the message itself — no preamble, no quotation marks, no Markdown fence. The user previews it before it is inserted, so Ask MUST be read-only: a draft that books the meeting it describes is a side effect nobody approved. Say so in your agent prompt, every time.

/imagine — a prompt becomes an image · optional

Finalizer fieldTypeMeaning
aspect_ratio"1:1" | "3:2" | "2:3", default "1:1" Unsupported values get bad_request.
quality"low" | "medium" | "high", default "high" Provider-relative.

Generating the same person, pet, or object across many prompts is backend behavior on top of this route, not a contract change: see repeatable Imagine references.

6 · Typed terminal results

The terminal result event wraps a discriminated result object; result.type is fixed per route. A client dispatches on the type, so a proxy that crosses wires fails loudly instead of inserting an image caption as a text message.

{"event": "result", "op_id": "op_7f3a…", "request_id": "req_5a09…",
 "audio": {"frames": 47, "bytes": 1504256, "duration_ms": 11200},
 "result": { …typed, see below… }}

audio echoes what the server actually consumed (absent for text input). The per-route result types:

// /dictate → "dictation"
{"type": "dictation",
 "text": "Let's push the review to Thursday afternoon.",
 "raw_transcript": "lets push the review to thursday afternoon",
 "polish_applied": true,
 "stt_route": "stream"}          // "stream" | "fallback", §8

// /ask → "message"
{"type": "message",
 "text": "Hey Sam, I'm running about ten minutes late.",
 "transcript": "tell sam i'm running ten minutes late"}   // audio input only

// /imagine → "image"
{"type": "image",
 "mime_type": "image/png",
 "byte_length": 184320,
 "sha256": "9f86d081884c7d65…",
 "data_base64": "iVBORw0KGgo…",
 "transcript": "a lighthouse at dusk in watercolour",     // audio input only
 "provider": "my-image-tool"}                             // optional

7 · Vocabulary injection

Recognizers mangle exactly the words users care most about — names, products, jargon. The vocabulary array in start is the client's per-operation term list, and a backend advertising "vocabulary": true MUST route it to wherever it can take effect:

RuleDetail
ShapeUnique strings, each 1–64 characters, at most 200 entries (or the advertised limit). Earlier entries are higher priority when the recognizer caps out.
SemanticsBiasing, not commands: an entry makes its surface form more likely, it never rewrites meaning.
PrivacyVocabulary is user content — a contact list in miniature. Do not log it, do not persist it beyond the operation, never expose it on /health.
Capability offAdvertise "vocabulary": false and clients omit the field. A backend MUST NOT advertise true and drop the list on the floor.

8 · Reliability: recovery, idempotency, fallback

V4 has one durability story and it is the client. The client already holds the audio (it recorded it); the backend keeps no state a lost connection strands. Three mechanisms make that safe:

9 · Errors and close codes

The terminal error event, on every route:

{"event": "error", "op_id": "op_7f3a…", "request_id": "req_2bda…",
 "code": "transcription_failed",
 "message": "speech recognition unavailable",
 "retryable": true}

message is shown to the user verbatim — keep it short and free of internals. The server closes with the mapped WebSocket code after sending the event. The code list is append-only; a client meeting an unknown code falls back to retryable, so set it truthfully — it is the only field a future client can rely on.

codeCloseMeaning
unauthorized4401Missing or invalid credential.
rate_limited4429Per-credential limit hit. Retryable.
not_supported4404Route not served, or text input where text_input is false. Check /health.
protocol_error4400Wrong first frame, malformed JSON, unknown control type, wrong protocol, audio before ready, or binary frames on text input.
bad_request4400Invalid start/finalize parameters, oversized frame, bad vocabulary shape.
timeout4408No start in 10 s, or a 60 s audio gap. Retryable.
audio_incomplete4409Finalize totals disagree with what the server received. Retryable — replay the operation.
audio_too_long4413Audio exceeded the advertised bound.
audio_too_short4422Finalize with under 300 ms of audio.
no_speech_detected4422Healthy transcription found no words.
transcription_failed4503Every transcription route failed. Retryable.
generation_failed4503The route's provider failed (agent for /ask, image provider for /imagine). Retryable.
internal_error4500Unexpected server error. Retryable.

Normal completion and cancel close with 1000.

10 · Retention and privacy

11 · Conformance

A backend conforms when it:

A client conforms when it:

A runnable conformance checker ships with the public reference implementation — see reference. Point it at your backend and it walks this checklist for you:

cd caret-docs/reference/go
go run ./cmd/caret-v4-conform -url https://your-backend.example -key "$CARET_API_KEY"