Your Agent · the caret/v4 protocol
Your keyboard, your agent, your machine.
Caret is a compose surface for iOS: speak or type an intent at the
cursor, preview, insert a finished message. In Your Agent mode the
keyboard talks to a backend you run yourself over the open
caret/v4 protocol: you enter a base URL and a credential,
and no third party sits in the middle of your typing. This site is
that protocol, plus two reference backends you can run today.
Three backend modes
| Mode in the app | What you enter | What it serves | Docs |
|---|---|---|---|
| Offline | Nothing | On-device speech recognition. No network, no backend, dictate only. | Nothing to configure |
| Caret API | Your Caret account | Caret's own hosted dictation, set up inside the app. Not a developer API. | Nothing to document |
| Your Agent | A base URL and the credential your backend expects | Whatever your backend advertises on /health: dictate always, ask and imagine when wired. |
Protocol, Reference |
The protocolstart here
One base URL serving three WebSocket routes — /dictate,
/ask, /imagine — over one shared live-audio
lifecycle: stream speech up, watch partial transcripts arrive while
you talk, send the route's finalizer, get exactly one typed terminal
result. Health and capability discovery on GET /health,
auth owned by the backend, per-operation vocabulary injection, and a
reliability story built on client-side replay. The protocol page is
the complete, normative contract — everything you need to implement
either side. Most people never need to: the reference backends
below implement it already, and you configure rather than write.
Have your agent set it upone URL
Give a coding agent
docs.typewithcaret.com/agent/instructions.md
and it clones the reference backend, wires its adapters to the
recognizer and model you already use, secures the credential, runs
the conformance checker, and does not call it done until a
sentence you dictated comes back through the app. No backend
written from scratch.
Any client, any backendno privileged party
The Caret keyboard is one client of this protocol, not its owner. Anyone may implement a backend — your agent, your model, your hardware, any language — and anyone may implement a client, and two conforming implementations interoperate without coordination. Two public reference backends ship in this repository, in Go and in Python, each with a conformance checker you can point at any V4 backend. Clone one, wire its lanes to your tools, and you have a backend; write your own only for a runtime they cannot run on.
The three routes
| Route | What happens | Terminal result |
|---|---|---|
/dictate · required |
Speech becomes polished, insert-ready text. Partial transcripts render while you speak; the final pass applies transcript cleanup unless you opt out. | "dictation" |
/ask · optional |
An instruction — spoken or typed — becomes one message, previewed before it is inserted. | "message" |
/imagine · optional |
A spoken or typed prompt becomes an image, delivered as verified bytes. | "image" |
Dictate is mandatory: a backend that cannot take speech is not a Caret
backend, and says so in /health rather than pretending.
Ask and Imagine are honestly optional — a dictate-only backend is a
complete backend, and the keyboard hides what a backend does not
advertise.
Why a live socket
- You see your words as you say them. Every route streams cumulative partial transcripts while the user speaks, so the text is on screen before the sentence ends.
- One lifecycle to implement. The routes differ only in their finalizer and their typed result; the connection handling, audio framing, reliability rules, and error table are identical across all three.
- Reliability without server state. The client keeps the audio it captured until a result is in hand; finalize totals prove the server heard everything; idempotent replay on a fresh connection recovers from any dropped transport without double work.
What the cleanup model is toldcaret-cleanup/1
Cleanup is the one place a backend hands the user's own words to a language model, so the wording is published rather than buried in source: the transcript travels as inert data, meaning is preserved completely, only formatting changes, and a failed cleanup returns the raw transcript. The spec is versioned independently of the protocol; V4 adds per-operation vocabulary injection into its glossary.
Design principles
- No privileged server. In Your Agent mode the
keyboard talks to the base URL and credential you enter. Any
server that implements
caret/v4works. - Capabilities are honest.
/healthdeclares what you actually serve, per route and per input mode; the keyboard hides the rest rather than failing in the user's hands. - Typed terminal results. Exactly one
resultorerrorper operation, tagged by route, after all partials — the last text the client holds is always authoritative. - Stable errors. An append-only code list with a
truthful
retryableon every error, mapped to WebSocket close codes. - Additive evolution. Unknown JSON fields are ignored on both sides; breaking changes take a new protocol number, not a silent shape change.
This documentation and the reference material are source-available, not open source: licensed for configuring, running, and extending Caret integrations only. See the license.