Troubleshooting
Common AI Transport problems and how to fix them. Each entry follows the same shape: the symptom you see, what to check to confirm the cause, and the fix to apply.
Entries are ordered roughly by how often they cause support tickets, most common first.
Channel namespace not configured for AI Transport
Symptom: Users see an empty assistant message, or a message containing only the first token. The server errors on subsequent publishes. The initial message.create succeeds, but the message.append operations that AI Transport uses to stream tokens fail with error code 93002 (Can only update/delete/append messages on channels with mutableMessages enabled) because the namespace does not permit appends. This is the single most common AI Transport failure.
Confirm:
- Check the server logs for error
93002after the first token. The message readsCan only update/delete/append messages on channels with mutableMessages enabled. - Open the Ably dashboard for your app's settings.
- Find the channel namespace your conversations live on (for example a namespace of
conversationsshould have channel names likeconversations:abc). - Check whether Message annotations, updates, deletes, and appends is enabled on that namespace.
Fix:
- Enable the Message annotations, updates, deletes, and appends rule on the namespace, using the dashboard, Control API, or CLI.
- Note that enabling this rule causes messages to be persisted regardless of whether persistence is enabled on the namespace.
Capability or token scope mismatch
Symptom: The token authenticates, but specific operations fail. There are two common shapes:
- Channel pattern mismatch: channel attachment fails, or a publish is rejected. The capability covers
conversations:*but your app useschat:abc. - Missing operation: the connection works but a specific operation does not. Clients cannot cancel generation (missing
publish). Late joiners or reconnects show only live messages and not the prior conversation (missinghistory). Agent presence never updates (missingpresence).
Confirm:
- Decode the JWT returned by your auth endpoint and inspect the
x-ably-capabilityclaim. It is a JSON-encoded map of channel patterns to permitted operations. - Compare each operation your application performs against the capability for the relevant channel name.
Fix:
- Make the channel pattern in the capability cover the channel names your application uses. Patterns are case-sensitive; wildcards like
conversations:*only cover channels with that exact prefix. - Grant every operation the application needs:
subscribeandpublishfor messages,historyfor loading past conversation,presencefor agent presence. Check the full list of capability operations against the AI Transport capability shape.
History disappears
Symptom: Past conversation messages are not available when the user expects them. Two scenarios trigger the same root cause:
- The user opens the app the next day, tries to scroll back, and older messages are gone.
- The user is offline longer than the live recovery window. On reconnect the SDK falls back to history, but no past messages arrive.
In both cases the cause is the same: the channel namespace is not configured to persist messages long enough for the use case.
Confirm:
- Check the retention setting on the channel namespace in the Ably dashboard. The default in-memory retention period covers only the live recovery window (around 2 minutes) and is not suitable for scroll-back.
- Confirm whether your application is meant to read history from Ably alone, or to hydrate from an external store.
Fix:
- Enable persistence on the channel namespace and set a retention period that covers your expected scroll-back window.
- For conversations that need to be retained for longer than the channel allows, persist completed turns to your own store and hydrate from it when the client reconnects.
Run never ends
Symptom: The streamed message renders with the streaming status forever. session.view.runs() shows the run with status: 'active' long after the model finished generating.
Confirm:
- Check the server logs around the affected run. Did
run.end({ reason })execute? If the route handler threw betweenrun.pipeandrun.end, the run never closes. - Inspect the channel in the Ably dashboard. The
ai-run-endlifecycle message should appear after the streamed message's close event. If it is missing,run.end()never ran or its publish failed.
Fix:
- Wrap streaming work in
try/finallyand always callrun.end()in thefinallyblock. A run that errored should end with reason'error'. - If you use Next.js
after(), confirm the callback runs to completion. An unhandled promise rejection insideafter()aborts the rest of the handler, includingrun.end(). - Check your lifecycle handling against the run lifecycle contract.
Cancel doesn't stop the agent
Symptom: The client publishes a cancel signal, the cancel message arrives on the channel, but the agent keeps streaming tokens until the model finishes naturally.
Confirm:
- Check whether the run's
abortSignalever fired. Tokens still arriving throughrun.pipemean it did not, and there are two reasons for that: anonCancelhook that returnedfalseand declined the cancel, or a cancel keyed to a run id the agent never matched. Output your agent publishes withrun.sendkeeps arriving even after the signal fires, because onlyrun.pipechecks the signal for you. - For long-running tools, check whether the tool implementation reads
run.abortSignal.abortedand exits when the signal fires. - If
onCancelis configured, check that it returnstruefor the cancel request. A hook that returnsfalserejects the cancel silently.
Fix:
- Pass
run.abortSignalto every LLM call, for examplestreamText({ abortSignal: run.abortSignal, ... }). - In server-executed tools, pass
run.abortSignalinto long-running operations so they exit promptly when the signal fires. - Trace the signal end to end against the cancellation flow.
Duplicate or unexpected turns
Symptom: A single user action produces two turns. The streamed response duplicates, or the user sees two siblings where they expected one. Two common causes:
- React Strict Mode (or a stale
useEffect) callssend()twice in development. - The user edits or regenerates a message while a previous turn is still streaming. The edit does not cancel the in-progress turn, so both streams run side by side.
Confirm:
- Inspect the channel for two
ai-run-startevents with differentrunIdvalues for the same user message. - Check whether your send path lives inside a
useEffect. Imperative event handlers (onClick,onSubmit) are safer.
Fix:
- Guard
view.send()so it fires once per user action. Avoid placing it inside an effect without a dependency that prevents re-firing. - Before editing or regenerating, cancel the in-flight run explicitly, following the recommended edit pattern. AI Transport does not auto-cancel an active run on edit.
Message too large to publish
Symptom: A publish fails with error code 40009 (maximum message length exceeded). Inside a streaming turn this surfaces as RunResponseStreamFailed (104008) on the server. Tool outputs or model responses that contain large payloads do not reach the channel.
Confirm:
- Identify the message that failed. Tool results that include binary data, large embeddings, or full document bodies are the usual culprits.
- Check your Ably package's message size limit. The cap is 64 KiB on Free and Standard, 256 KiB on Pro and Enterprise.
Fix:
- Stream large tool results across multiple events instead of publishing them in one message.
- Persist large payloads to an external store and send only a reference (URL or ID) over the channel.
Two devices share a clientId
Symptom: Ownership-scoped behaviour misbehaves across two devices for the same user. An onCancel hook that authorises by clientId accepts either device's cancel, so one device stops the other's turn. Presence shows the user appearing and disappearing as both devices update their state.
Confirm:
- Inspect the JWT each device receives. If both tokens have the same
x-ably-clientId, they are indistinguishable to the Ably service. - Check whether your token-issuing logic generates a unique identifier per device (typically
userId + deviceId), or only usesuserId.
Fix:
- Assign a unique
clientIdper device or per session for any case where ownership matters. A common pattern is<userId>:<deviceId>so the user remains identifiable while devices remain distinguishable. - Check your ownership assumptions against the multi-device session model.
Branch selection out of sync across devices
Symptom: Two devices on the same conversation show different responses to the same prompt. Edit history navigation diverges between them.
The divergence is intentional. Branch selection is a per-view, per-device choice. Each device navigates the conversation tree independently, so a user reviewing alternative replies on one device does not disrupt another device displaying the same conversation.
Fix:
- If branch selection should be shared (for example, a co-pilot where both devices must agree), synchronise the conversation tree's selected leaf through your own application state, or by sending a message on the channel that the other client can react to and change its selections.
- Confirm how branch selection works before you design the sync.
Reconnect loop
Symptom: The client connects, drops, reconnects, drops again. The cycle repeats every few seconds. Channel state never stabilises.
Confirm:
- Inspect the Ably connection state on the client (
realtimeClient.connection.state). A reconnect loop cyclesconnecting→connected→disconnectedrapidly. - If each
disconnectedevent carries a token error, your auth endpoint is returning tokens that Ably is rejecting on every reconnect.
Fix:
- Verify the auth endpoint signs tokens with the correct key for the right environment (a dev key against a prod app is a common cause).
- Lengthen token TTL if it is set very short. The SDK refreshes ahead of expiry; very short lifetimes fight the refresh.
- If the token is rejected for capability rather than signing, see capability or token scope mismatch.
Agent process crashes mid-stream
Symptom: A streamed message stops part-way through and stays in streaming status forever. No more tokens arrive and no run-end event ever fires. From the client's point of view this is indistinguishable from Run never ends: both leave the run in active state. The root cause here is the process dying rather than the handler skipping run.end().
Confirm:
- Check server logs for the crashed process around the timestamp of the affected turn. An exception or an OOM kill typically appears there.
- If you cannot see the crash directly, check infrastructure signals: serverless function timeouts that expire before
after()finishes, container OOM kills, and process restarts are common causes.
Fix:
- Add structured error logging around the streaming path so the cause is visible. Common causes are model provider errors, OOM in long tool calls, and serverless function timeouts.
- Decide on the application response: surface a retry control to the user, or auto-retry from the application layer. AI Transport does not retry the LLM call automatically.
Publish fails in suspended state
Symptom: A publish from the client returns an error. The connection has been disconnected for more than around two minutes and has moved to suspended state.
Confirm:
- Inspect the connection state at the moment of the failed publish. A
suspendedconnection has lost message continuity, and there is no live connection to the Ably platform. The session checks the channel state before it publishes, and throwsSessionChannelNotReady(104007) with the messageunable to send; channel is suspended.
Fix:
- Check connection state before publishing user-driven actions, and queue them locally while the connection is not
connected. - Flush the queue once the connection returns to
connected.ably-jsdoes not buffer publishes through asuspendedstate because continuity has already been lost.
The client never learns the run id
Symptom: send() resolves, the user's message appears, and the agent streams a reply, but run.started never settles and run.runId stays empty. A Stop button gated on the run id is never enabled.
Confirm:
- Inspect the channel in the Ably dashboard. Find the
ai-run-startthe agent published and read itsinput-codec-message-idheader. It has to match thecodec-message-idon the input the client published. - Check what the agent passed to
createRun. The run is keyed on the input the invocation refers to, so an agent that builds its own invocation, or reuses one from an earlier turn, opens a run against a different input. - Check that the POST that woke the agent carried this send's
run.toInvocation().toJSON()rather than a value your application constructed.
Fix:
- Pass the invocation straight through from
run.toInvocation()to the agent, and rebuild it there withInvocation.fromJSON. Nothing else identifies which input the run answers. - A run that legitimately has no client waiting on it, such as an agent-initiated turn, publishes no
input-codec-message-id. Do not gate your UI on a run id in that case; read the run status fromruns()instead.
A steering message rejects without publishing
Symptom: activeRun.steer(input) rejects both published and outcome immediately, and nothing appears on the channel.
Confirm:
- Read the run's status from
runs(). Steering is only accepted while the run is live. A run whose end the client has already observed rejects the call rather than publishing a message no agent will read. - Check whether a cancel or an error terminated the run just before the steering message went out.
Fix:
- Gate the steering control on
status === 'active'and drop it as the run ends. - If the user's input still matters, send it as a new turn instead. Interruption and steering covers the choice.
Published messages are ignored
Symptom: Something publishes to the conversation channel and nothing reaches the view. The message is visible in the Ably dashboard and no error is raised.
Confirm:
- Inspect the message's
extrasin the dashboard. Every message AI Transport decodes carries anextras.aienvelope holding the transport and codec headers. A message without one is foreign to the transport and is skipped. - Turn the SDK's log level up to debug. A foreign message is logged as ignored rather than raised as an error, because sharing a channel with application traffic is expected.
- Check what published it. A plain
channel.publish()from application code, or a Pub/Sub SDK client sharing the channel, produces messages in this shape.
Fix:
- Publish conversation content through the session or the codec's input factories, which add the
extras.aienvelope. - Keep application-level traffic off the conversation channel, or carry it in
extras.headers, which is reserved for the application and left untouched by the transport.
When to contact support
If you have worked through the entries on this page and the symptom does not match, or the fix did not work, capture:
- The channel name.
- The date and time of the affected operation.
- The
clientIdof the affected server agent or client device. - The first error log entry on either side.
Open a support ticket with these. The Ably side of the system is observable to support; the application side needs the IDs to correlate.