# AI Transport pricing How AI Transport operations contribute to your message count and strategies to optimize costs. The [AI Transport SDK](https://ably.com/docs/ai-transport.md) is built on top of [Pub/Sub](https://ably.com/docs/pub-sub.md). All AI Transport operations generate Pub/Sub messages that follow the same [counting rules as Pub/Sub](https://ably.com/docs/platform/pricing/message-counting.md), and Ably counts messages rather than tokens. Every message counts once when it is published, then once more for each subscriber that receives it. On an AI Transport session the subscribers are every user client attached to the session, every agent attached to the session, and the history store, because AI Transport [stores every message](#persistence). A publisher is one of those subscribers too: AI Transport relies on [`echoMessages`](https://ably.com/docs/pub-sub/advanced.md#echo) being enabled, so a client or agent receives its own message back. A message published on a session with one browser client and one agent therefore counts four times: once to publish it, once for the client, once for the agent, and once to store it. ## AI Transport operations The following table shows how AI Transport operations contribute to your message count: | Operation | Messages counted | | --- | --- | | [User input](https://ably.com/docs/ai-transport/concepts/runs.md#invocations) || | Send a message | 1 inbound message per message part | | Cancel a run | 1 inbound message | | Input delivery | 1 outbound message per subscriber | | [Token streaming](https://ably.com/docs/ai-transport/features/token-streaming.md) || | Open a streamed message | 1 inbound message | | Token append | 1 inbound message per published append after [rollup](#rollup) | | Close a streamed message | 1 inbound message | | Streamed message delivery | 1 outbound message per subscriber for the open, for each published append, and for the close | | [Run and step lifecycle](https://ably.com/docs/ai-transport/concepts/runs.md) || | Run start | 1 inbound message | | Run suspend | 1 inbound message | | Run resume | 1 inbound message | | Run end | 1 inbound message | | Step start | 1 inbound message | | Step end | 1 inbound message | | [Codec lifecycle event](https://ably.com/docs/ai-transport/internals/codec-architecture.md#two-layer-split) | 1 inbound message per event | | Lifecycle event delivery | 1 outbound message per subscriber per event | | [History and replay](https://ably.com/docs/ai-transport/features/history.md) || | History retrieval | 1 outbound message per retrieved message | | [Agent presence](https://ably.com/docs/ai-transport/features/agent-presence.md) || | Presence enter | 1 inbound message | | Presence leave | 1 inbound message | | Presence update | 1 inbound message | | Presence event delivery | 1 outbound message per presence subscriber | Parts come from the message format your AI framework already uses. For example, with the Vercel AI SDK a message you send is a list of parts, so typing a prompt gives you one text part and attaching an image adds a second. AI Transport publishes each part as its own message, so a plain text prompt counts as one message and the same prompt with an image counts as two. Parts with nothing to send, such as the AI SDK's step markers, are skipped. Each streamed part of the agent response is also an independent message. For example, a response can open and close one message for reasoning and a second message for text output. Every tool call opens another. The [wire protocol](https://ably.com/docs/ai-transport/internals/wire-protocol.md#event-names) lists the message `name` for each type of operation or message published. AI Transport publishes `ai-run-start` and `ai-run-end` for the run, and `ai-step-start` and `ai-step-end` around each `run.pipe` call. On top of those, a codec can publish events of its own. The Vercel codec does: `start` and `finish` once per response, and `start-step` and `finish-step` once per model call, which is four extra messages for a plain text reply. The OpenAI codec publishes none, because AI Transport already tracks the run. Individual appends each count as one message, but Ably stores and serves a streamed response as one aggregated message, so retrieving it from history counts as one outbound message rather than one per append. Editing a message or regenerating a response starts a new run and emits the same set of events as any other run. The original [branch](https://ably.com/docs/ai-transport/features/branching.md) stays on the session. ## Append rollup An agent appends each token event to the streamed message as the model produces it, and Ably coalesces the appends that fall inside the rollup window into a single published message. Each published message is what gets counted, inbound and then outbound to every subscriber, so the rollup window sets the ceiling on the message count for a stream. The default window of 40ms caps a single response at 25 published messages per second, or the model's output rate if that is lower. Increasing the window length lowers the number of messages per second: at 100ms a stream publishes at most 10 messages per second. Clients receive the same content either way, delivered in fewer and larger updates. The [rollup configuration table](https://ably.com/docs/ai-transport/features/token-streaming.md#configure-rollup) provides more information on setting the `appendRollupWindow` value. ## Sessions and channels Each AI Transport [session](https://ably.com/docs/ai-transport/concepts/sessions.md#session-and-channel) maps to a single Ably channel. Everything in the conversation shares that channel: user input, streamed responses, lifecycle events, presence, and LiveObjects state. [Concurrent turns](https://ably.com/docs/ai-transport/features/concurrent-turns.md) are multiplexed onto the same channel rather than opening more. Channel time contributes to your [channel minutes](https://ably.com/docs/platform/pricing.md#channels). A channel becomes inactive approximately one minute after the last subscriber detached or the last message was published. The billable channel time for a conversation therefore tracks how long a client or an agent stays attached rather than how many messages pass through it. ## Connections Ably bills each connected client for [connection minutes](https://ably.com/docs/platform/pricing.md#connections). A connection-minute is counted for every minute a client maintains an open connection, regardless of activity. Clients that remain connected but idle still accrue connection minutes. An agent holds a connection too. Closing a session detaches its channel but leaves the Ably client connected, so connection minutes stop only when the application closes the client or the process exits. A serverless agent that exits after each response accrues connection minutes only while it runs; a long-lived agent process accrues them continuously. ## Message persistence AI Transport requires the **Message annotations, updates, deletes, and appends** [rule](https://ably.com/docs/ai-transport/getting-started/channel-rules.md) on the namespace your sessions live on, because appends only work on channels that have this rule enabled. The rule also stores every message on those channels, which is why each message counts once for the history store. This applies to the entire namespace of channels referenced by the rule, not just to those channels carrying AI Transport sessions. ## Cost optimization ### Tune the append rollup window Raise [`appendRollupWindow`](https://ably.com/docs/ai-transport/features/token-streaming.md#configure-rollup) above the 40ms default on the connection that publishes the stream. This is the largest single lever on the message count of a streamed response, because it caps how many appends Ably publishes per second. Trade it against how granular you want token delivery to feel. ### Only enable persistence for channels that require it Enable the rule on a namespace that carries conversations only, such as `conversations:`, rather than on a broad prefix. The rule stores every message on every channel in that namespace, so a narrow namespace keeps unrelated channels out of your stored message count. ### Hydrate long conversations from your own store Loading history from the channel counts one outbound message per retrieved message. Use [database hydration](https://ably.com/docs/ai-transport/features/database-hydration.md) to serve the stored part of a conversation from your own database and let the session cover only what is live. ### Close sessions and connections when the conversation stops On the client, call `session.close()` when the user navigates away. On the agent, call `session.end()` when the response completes, or `session.detach()` when a durable activity is handing an open run to the next one. Each of these detaches the channel so it can become inactive, which means that there are no further active channel minutes. Closing the Ably client is what stops connection minutes, on both sides. ### Keep steps proportionate to the work Each [step](https://ably.com/docs/ai-transport/concepts/runs.md#steps) publishes two lifecycle events, and each `run.pipe` call brackets a step of its own. Bracket any further steps around units you want to retry independently rather than around every model call. ## Read next - [Ably pricing](https://ably.com/docs/platform/pricing.md): the current message, channel, and connection rates. [Contact Ably](https://ably.com/contact) for a quote. - [AI chatbot pricing example](https://ably.com/docs/platform/pricing/examples/ai-chatbot.md): a worked monthly cost for a chatbot handling 35,000 conversations. - [Going to production](https://ably.com/docs/ai-transport/going-to-production.md): the limits, retention, and monitoring to settle alongside cost. ## Documentation Index To discover additional Ably documentation: 1. Fetch [llms.txt](https://ably.com/llms.txt) for the canonical list of available pages. 2. Identify relevant URLs from that index. 3. Fetch target pages as needed. Avoid using assumed or outdated documentation paths.