Reconnection and recovery

Your users get the whole response, even when their connection drops. The agent keeps publishing to the channel, and a client that drops resumes from the point it reached.

Streams survive connection drops. The agent carries on publishing to the channel whether or not a client is attached, and a client that dropped picks up from the serial it reached when it reconnects.

Diagram showing a client disconnecting mid-stream and resuming from the exact point of disconnection on reconnect

How it works

The channel outlives every connection attached to it. When a client's connection drops:

  1. The agent keeps publishing. Its output goes to the channel rather than down that client's connection, so the run carries on unchanged.
  2. The Ably client reconnects on its own.
  3. On reconnect, delivery resumes from the last serial the client received, so messages published during the gap arrive.
  4. For a gap longer than the live recovery window, your code pages history instead of relying on replay.

Recovery scenarios

The recovery path depends on how long the client was disconnected.

When the disconnection is brief, Ably's connection protocol handles it. Delivery resumes from the last serial the client received, and messages published during the gap arrive in order, so the response continues where it stopped.

When the client has been offline longer than the live recovery window, continuity is lost and it has to read what it missed. That is a history walk backwards from the new attach point, and it is the one recovery step your code handles:

JavaScript

1

2

3

// Client-side, after a reconnect that lost continuity.
const { events } = await transport.history({ limit: 100 });
for (const event of events) fold.apply(event);

A replayed wire message is deduped on its append version serial, so merging a message you already had does not double it.

Server-side encoder recovery

On the agent side, the encoder handles transient failures during streaming. If an append fails, for example on a network blip between your agent and Ably, the encoder falls back to a full message update:

  1. Append the next token to the message (normal path).
  2. If the append fails, send a full update with the accumulated content (recovery path).
  3. The encoder sends that update once, when it closes the stream, carrying the terminal status.

This happens inside run.pipe(). Subscribers still receive the whole accumulated response, even when individual append operations fail.

Mid-stream joins

A client that attaches while a response is already streaming never saw that stream's start. The codec's lifecycle tracker synthesises the stream-start events it missed, so the decoded stream arrives well-formed and your reducer handles it like any other.

A second tab opened mid-answer therefore shows the answer so far, then the remaining appends live.

Know when to reach for history

Most reconnects need nothing from you. Reach for history in two cases:

  • The connection lost continuity, so replay could not cover the gap. Ably surfaces this on the connection state.
  • The client is attaching for the first time, or after a reload, and has nothing to render.

Everything else, including a brief drop mid-stream, is handled by the connection resuming from its last serial.

Edge cases and unhappy paths

  • A client that drops mid-stream and reconnects after the live recovery window receives the accumulated content of the message up to the latest append rather than a replay of every individual token. The user-visible result is the same.
  • A client without channel history capability cannot reconnect after the live recovery window. Capability scoping is part of authentication.
  • An agent that crashes mid-stream leaves the partial message on the channel with its status header still streaming, because nothing closed it, and the run stays open. A fresh process that opens a new run publishes beside it rather than continuing it; a durable-execution retry under the same stepId supersedes it instead.
  • The encoder fallback to a full message update is invisible to subscribers. If you log channel operations, you see one update per stream that had a failed append, published when the encoder closes it.
  • A client clock drift does not affect recovery. Reconnection uses the channel's serial rather than wall-clock time.

FAQ

What is the live recovery window?

It is the period during which Ably can replay messages without falling back to history. The duration depends on the connection state and the channel configuration. After that window, your code pages history to fill the gap.

Does the user see the agent pause when they reconnect?

No. The client receives the accumulated content of the streamed message, then the remaining appends. Rendering stays continuous as long as your reducer keys accumulated text by meta.codecMessageId rather than appending blindly.

How long is the response retained after the agent ends the turn?

The channel keeps it for the period configured by your retention window. Configure that window through the channel's persistence settings, then recover past responses with history and replay.

What if the agent process dies before the stream finishes?

The partial message stays on the channel with its status header at streaming, because nothing closed it, and the run stays open. The conversation is intact. AI Transport does not retry the model call for you. To survive an agent-process crash mid-run without abandoning the turn, run the agent inside a durable workflow engine and adopt the run from the retry process.

Does the client need special code to handle reconnection?

Almost none. Reconnection and replay are the Ably connection's own behaviour. The one case your code handles is a gap too long for replay, where you page history to fill it.