Caret Docs typewithcaret.com →

Your Agent · the caret/v4 protocol

Your keyboard, your agent, your machine.

Caret is a compose surface for iOS: speak or type an intent at the cursor, preview, insert a finished message. In Your Agent mode the keyboard talks to a backend you run yourself over the open caret/v4 protocol: you enter a base URL and a credential, and no third party sits in the middle of your typing. This site is that protocol, plus two reference backends you can run today.

Three backend modes

Mode in the appWhat you enterWhat it servesDocs
Offline Nothing On-device speech recognition. No network, no backend, dictate only. Nothing to configure
Caret API Your Caret account Caret's own hosted dictation, set up inside the app. Not a developer API. Nothing to document
Your Agent A base URL and the credential your backend expects Whatever your backend advertises on /health: dictate always, ask and imagine when wired. Protocol, Reference

The protocolstart here

One base URL serving three WebSocket routes — /dictate, /ask, /imagine — over one shared live-audio lifecycle: stream speech up, watch partial transcripts arrive while you talk, send the route's finalizer, get exactly one typed terminal result. Health and capability discovery on GET /health, auth owned by the backend, per-operation vocabulary injection, and a reliability story built on client-side replay. The protocol page is the complete, normative contract — everything you need to implement either side. Most people never need to: the reference backends below implement it already, and you configure rather than write.

Read the caret/v4 protocol →

Have your agent set it upone URL

Give a coding agent docs.typewithcaret.com/agent/instructions.md and it clones the reference backend, wires its adapters to the recognizer and model you already use, secures the credential, runs the conformance checker, and does not call it done until a sentence you dictated comes back through the app. No backend written from scratch.

Agent instructions →

Any client, any backendno privileged party

The Caret keyboard is one client of this protocol, not its owner. Anyone may implement a backend — your agent, your model, your hardware, any language — and anyone may implement a client, and two conforming implementations interoperate without coordination. Two public reference backends ship in this repository, in Go and in Python, each with a conformance checker you can point at any V4 backend. Clone one, wire its lanes to your tools, and you have a backend; write your own only for a runtime they cannot run on.

The reference implementation →

The three routes

RouteWhat happensTerminal result
/dictate · required Speech becomes polished, insert-ready text. Partial transcripts render while you speak; the final pass applies transcript cleanup unless you opt out. "dictation"
/ask · optional An instruction — spoken or typed — becomes one message, previewed before it is inserted. "message"
/imagine · optional A spoken or typed prompt becomes an image, delivered as verified bytes. "image"

Dictate is mandatory: a backend that cannot take speech is not a Caret backend, and says so in /health rather than pretending. Ask and Imagine are honestly optional — a dictate-only backend is a complete backend, and the keyboard hides what a backend does not advertise.

Why a live socket

What the cleanup model is toldcaret-cleanup/1

Cleanup is the one place a backend hands the user's own words to a language model, so the wording is published rather than buried in source: the transcript travels as inert data, meaning is preserved completely, only formatting changes, and a failed cleanup returns the raw transcript. The spec is versioned independently of the protocol; V4 adds per-operation vocabulary injection into its glossary.

Read the cleanup spec → · prompt.md

Design principles

This documentation and the reference material are source-available, not open source: licensed for configuring, running, and extending Caret integrations only. See the license.