AI Transport
AI Transport helps you build AI applications that are resilient, resumable, multi-user and multi-device. Adopt the streaming APIs and keep your own conversation store, or use durable sessions and let AI Transport hold the conversation for you.
AI Transport helps you build AI applications that are resilient, resumable, multi-user and multi-device. Agents and clients attach to an Ably channel that outlives any single connection, so the conversation survives a network drop and reaches all of the user's devices.
Streaming delivers your agent's responses to every subscribed client, in realtime. A client that disconnects can automatically reconnect and resume where it left off. Any subscribed client can interrupt, steer, or cancel the agent while it is running. You own the message format between the agent and the client, and you choose whether the conversation history lives in your own database or is read back from the channel.
Durable sessions provide the complete, persistent state of a conversation, and it exists independently of anything that connects to it. The conversation history can be fully hydrated from the Ably channel. The session rebuilds a conversation for any client that connects, updates the messages on screen as new messages arrive, and supports edits, regenerating, branching, and tool calls.
The AI Transport SDKs and the Ably channel provide the message delivery infrastructure. Your model provider, prompts, orchestration, and hosting remain the same. AI Transport ships a drop-in transport for the Vercel AI SDK and a codec for the OpenAI Responses API. For a provider with no bundled codec, you write a wire codec.
What streaming fixes
Most AI frameworks send one HTTP request per turn and return the model's output on that request's response body. This ties the response to the HTTP request, so other clients do not receive the response, and if the HTTP connection drops then the response is lost.
| Scenario | Direct HTTP streaming | With AI Transport |
|---|---|---|
| The connection drops mid-response | The response is lost, the model keeps producing output that is discarded, and the user sees an error and starts again. | The agent keeps publishing to the channel. The client reconnects and resumes from where it left off. |
| The user opens the conversation on a second device | The response belongs to the device that opened the request, so a second device cannot read it. | Every device subscribes to the same channel and reads the response as it streams. |
| The user reloads the page mid-response | The reload discards the request and the partial response with it. | The client resubscribes and reads the in-flight response back from channel history, where the response is a single message that grows as the model produces it. |
| The user presses stop | Closing the connection is the only signal available, and the agent cannot tell it apart from a network drop. | The client publishes a cancel signal. The agent's current run ends and the channel stays open. |
| A second message arrives before the first response finishes | A second request opens a second stream, and the client interleaves two responses or drops one. | Each turn is its own run on the same channel, with its own output and its own cancel signal. |
| The agent process restarts mid-response | The outbound stream ends with the process, and the client is left displaying a truncated response. | With durable execution, the retried step publishes to the same channel and supersedes the output of the attempt that failed. |
To solve these yourself, you need a buffer so a stream can resume, a store for conversation state, and a queue or a second WebSocket for the client to signal on. HTTP streaming and AI goes through each limitation and the workaround code it requires.
How AI Transport is built
The Ably channel
Every conversation is one Ably channel: a durable, append-only log that any client or agent attaches to by name. Messages outlive the connection, device, or process that published them, and they are stored in order for the period configured by your retention window. A client that drops can resubscribe and read the rest of the conversation without gaps or duplicate messages. Message order is guaranteed from a single publisher to a single region, and message ordering covers what order means when publishing in different regions.
AI Transport operates on a standard Ably channel, so you can make use of other Ably channel features. Agent presence can be used to report if an agent is thinking, streaming, idle, or offline, and lets an agent stop work when no client is connected to read the response. LiveObjects state gives the client and the agent one shared store to read and write. Push notifications reach a user whose client is not connected.
The streaming layer
Streaming makes it easy to deliver messages to clients at scale. A client transport publishes the user's input and reads the agent's output, and an agent transport opens a run and pipes the model's output into it, allowing the channel to deliver that output to any subscribing client. A codec maps your framework's event types onto channel messages, and the model's output streams by appending to a single message, so a client that subscribes late reads one aggregated response instead of each token individually. A run brackets one turn of agent work, and a step is a retryable unit inside a run, where a later attempt under the same identifier supersedes the earlier one, which is how a retried step replaces failed output already on the channel. Any client or agent can publish, which allows a client to cancel and steer an agent in realtime.
Durable sessions
Durable sessions build on top of the streaming layer, adding in conversation storage, management and advanced features. Client and agent sessions own the channel attachment, the run lifecycle, and cancel routing. The conversation tree makes the user's prompt and the agent's reply each a tree node with a parent and siblings, and a view selects one path through the tree's branches as an ordered list your UI can easily make use of. React hooks bind a view to a component: useView returns the messages, the run statuses, and the write operations, and re-renders when the tree changes underneath it.
What durable sessions add
Each of the following comes with the session rather than being something you model yourself:
| What you get | How it works |
|---|---|
| Branching, edit, and regenerate | Editing a message or regenerating a response forks the conversation instead of overwriting it, and the tree keeps every branch. |
| The channel as the conversation storage | A client that connects rebuilds the whole conversation from the channel. |
| Tool calling | A tool call and its result are part of message data, so every subscribed client sees the call and the answer. |
| Human-in-the-loop approval | A run suspends until a tool call is approved. The request waits in the session, so any user can answer it from any device, minutes or hours later. |
| Database hydration | History you have persisted yourself can be inserted into the live conversation with no gaps and no duplicate messages. |
Keep the stack you have
In the agent, you publish the stream you already build into a run on the channel rather than returning it as the HTTP response body. Prompts, tool definitions, model calls, and rendering are untouched. AI Transport implements the Vercel AI SDK's ChatTransport interface, so useChat accepts it as its transport.
Streaming and durable sessions each show a client and an agent in full.
Limits and dependencies
These apply to every application built on AI Transport:
| Limit or dependency | Detail |
|---|---|
| Your retention window bounds what the channel holds | The channel holds the conversation for as long as channel retention covers it, a window you set per channel rule, so the channel is always the live-recovery window. Keep your own store of completed runs for history beyond that window, and database hydration joins the stored history to the live conversation. |
| Hydrated history is linear | Messages you have persisted yourself cannot be edited or regenerated, so branching applies to the part of the conversation the channel still holds. |
| Channel content is visible to subscribed clients | Every message reaches every subscribed client whose token capability allows it, tool inputs and outputs included. Scope capabilities to channel name prefixes so a token only grants access to the conversations that client should read. |
| A network hop and a third-party dependency | Conversations go through Ably rather than your own infrastructure, so every publish makes a round trip. Append rollup batches the model's output, so that round trip happens once per batch. |
Going to production covers limits, retention windows, monitoring, auth hardening, and pricing.
Platform guarantees
Ordering, persistence, replication, and regional failover are guarantees of the Ably platform, which is designed for 99.999% global service availability, delivers at low latency across regions, and scales elastically. They apply to AI Transport in the same way as to every other Ably product.
Ably is SOC 2 Type II certified and HIPAA compliant, and operates a bug bounty program.
Pick where to start
Streaming
Stream your agent's responses to every subscribed client, and let any of them steer it.
Durable sessions
Let AI Transport hold the conversation, and rebuild it for any client that connects.
Durable execution
Keep an agent turn safe across a process restart, alongside either tier.
Going to production
The checklist to work through before you ship.
Read next
- HTTP streaming and AI: each limitation of direct HTTP streaming, and the workaround code it requires.
- Get started with streaming: a working chat app whose conversation stays in your own store.
- Get started with durable sessions: the same app, with AI Transport holding the conversation.