AI Transport pricing

How AI Transport operations contribute to your message count and strategies to optimize costs.

The AI Transport SDK is built on top of Pub/Sub. All AI Transport operations generate Pub/Sub messages that follow the same counting rules as Pub/Sub, and Ably counts messages rather than tokens.

Every message counts once when it is published, then once more for each subscriber that receives it. On an AI Transport session the subscribers are every user client attached to the session, every agent attached to the session, and the history store, because AI Transport stores every message. A publisher is one of those subscribers too: AI Transport relies on echoMessages being enabled, so a client or agent receives its own message back.

A message published on a session with one browser client and one agent therefore counts four times: once to publish it, once for the client, once for the agent, and once to store it.

AI Transport operations

The following table shows how AI Transport operations contribute to your message count:

OperationMessages counted
User input
Send a message1 inbound message per message part
Cancel a run1 inbound message
Input delivery1 outbound message per subscriber
Token streaming
Open a streamed message1 inbound message
Token append1 inbound message per published append after rollup
Close a streamed message1 inbound message
Streamed message delivery1 outbound message per subscriber for the open, for each published append, and for the close
Run and step lifecycle
Run start1 inbound message
Run suspend1 inbound message
Run resume1 inbound message
Run end1 inbound message
Step start1 inbound message
Step end1 inbound message
Codec lifecycle event1 inbound message per event
Lifecycle event delivery1 outbound message per subscriber per event
History and replay
History retrieval1 outbound message per retrieved message
Agent presence
Presence enter1 inbound message
Presence leave1 inbound message
Presence update1 inbound message
Presence event delivery1 outbound message per presence subscriber

Parts come from the message format your AI framework already uses. For example, with the Vercel AI SDK a message you send is a list of parts, so typing a prompt gives you one text part and attaching an image adds a second. AI Transport publishes each part as its own message, so a plain text prompt counts as one message and the same prompt with an image counts as two. Parts with nothing to send, such as the AI SDK's step markers, are skipped.

Each streamed part of the agent response is also an independent message. For example, a response can open and close one message for reasoning and a second message for text output. Every tool call opens another. The wire protocol lists the message name for each type of operation or message published.

AI Transport publishes ai-run-start and ai-run-end for the run, and ai-step-start and ai-step-end around each run.pipe call. On top of those, a codec can publish events of its own. The Vercel codec does: start and finish once per response, and start-step and finish-step once per model call, which is four extra messages for a plain text reply. The OpenAI codec publishes none, because AI Transport already tracks the run.

Individual appends each count as one message, but Ably stores and serves a streamed response as one aggregated message, so retrieving it from history counts as one outbound message rather than one per append.

Editing a message or regenerating a response starts a new run and emits the same set of events as any other run. The original branch stays on the session.

Append rollup

An agent appends each token event to the streamed message as the model produces it, and Ably coalesces the appends that fall inside the rollup window into a single published message. Each published message is what gets counted, inbound and then outbound to every subscriber, so the rollup window sets the ceiling on the message count for a stream.

The default window of 40ms caps a single response at 25 published messages per second, or the model's output rate if that is lower. Increasing the window length lowers the number of messages per second: at 100ms a stream publishes at most 10 messages per second. Clients receive the same content either way, delivered in fewer and larger updates. The rollup configuration table provides more information on setting the appendRollupWindow value.

Sessions and channels

Each AI Transport session maps to a single Ably channel. Everything in the conversation shares that channel: user input, streamed responses, lifecycle events, presence, and LiveObjects state. Concurrent turns are multiplexed onto the same channel rather than opening more.

Channel time contributes to your channel minutes. A channel becomes inactive approximately one minute after the last subscriber detached or the last message was published. The billable channel time for a conversation therefore tracks how long a client or an agent stays attached rather than how many messages pass through it.

Connections

Ably bills each connected client for connection minutes. A connection-minute is counted for every minute a client maintains an open connection, regardless of activity. Clients that remain connected but idle still accrue connection minutes.

An agent holds a connection too. Closing a session detaches its channel but leaves the Ably client connected, so connection minutes stop only when the application closes the client or the process exits. A serverless agent that exits after each response accrues connection minutes only while it runs; a long-lived agent process accrues them continuously.

Message persistence

AI Transport requires the Message annotations, updates, deletes, and appends rule on the namespace your sessions live on, because appends only work on channels that have this rule enabled. The rule also stores every message on those channels, which is why each message counts once for the history store.

This applies to the entire namespace of channels referenced by the rule, not just to those channels carrying AI Transport sessions.

Cost optimization

Tune the append rollup window

Raise appendRollupWindow above the 40ms default on the connection that publishes the stream. This is the largest single lever on the message count of a streamed response, because it caps how many appends Ably publishes per second. Trade it against how granular you want token delivery to feel.

Only enable persistence for channels that require it

Enable the rule on a namespace that carries conversations only, such as conversations:, rather than on a broad prefix. The rule stores every message on every channel in that namespace, so a narrow namespace keeps unrelated channels out of your stored message count.

Hydrate long conversations from your own store

Loading history from the channel counts one outbound message per retrieved message. Use database hydration to serve the stored part of a conversation from your own database and let the session cover only what is live.

Close sessions and connections when the conversation stops

On the client, call session.close() when the user navigates away. On the agent, call session.end() when the response completes, or session.detach() when a durable activity is handing an open run to the next one. Each of these detaches the channel so it can become inactive, which means that there are no further active channel minutes. Closing the Ably client is what stops connection minutes, on both sides.

Keep steps proportionate to the work

Each step publishes two lifecycle events, and each run.pipe call brackets a step of its own. Bracket any further steps around units you want to retry independently rather than around every model call.