OpenAI Responses
The OpenAI Responses API streams model output as typed events over HTTP. AI Transport's ResponsesCodec encodes that stream onto an Ably channel, so the same server code feeds durable sessions instead of an ephemeral HTTP response.
The OpenAI Responses API is OpenAI's interface for calling a model and getting back a stream of typed events: output text, reasoning, refusals, function-call arguments, and lifecycle events. AI Transport writes that stream to an Ably channel, so the channel gives you reconnects, multi-device delivery, and bidirectional control.
The OpenAI codecs in @ably/ai-transport/openai publish the Responses events a client needs to the channel in OpenAI's own event vocabulary and reassemble the conversation on the client. ResponsesCodec is the wire half, the encoder and decoder; ResponsesSessionCodec adds the reducer that rebuilds messages, and it is the one a session takes. There is no OpenAI-specific transport, so you use the generic createAgentSession from @ably/ai-transport and pass it ResponsesSessionCodec.
To build a working app with this codec, see Get started with OpenAI.
What the OpenAI Responses API brings
The openai npm package owns the model call and the typed event model:
| Feature | Description |
|---|---|
| Responses streaming | client.responses.create({ stream: true }) returns a stream of ResponseStreamEvents: output-text deltas, reasoning, refusals, function-call arguments, and item lifecycle events. |
| Server-side tools | You advertise function tools on the request. The model emits function calls, you run them, and you feed the outputs back on the next Responses API call. |
| Reasoning models | Reasoning items carry the model's summarised thinking, and encrypted_content for the no-store, zero-data-retention round trip. |
| Round-trippable items | Each output item the codec stores is also valid Responses API input, so the conversation feeds the next turn without conversion. |
| Typed SDK | The official SDK owns auth, the HTTP call, streaming, and the Responses type definitions. |
What AI Transport adds
Everything below comes from AI Transport rather than the framework.
| Feature | Description | Available with |
|---|---|---|
| Streams that outlive a connection | Tokens flow through a channel rather than one HTTP response, so a client can reconnect and pick up where it left off. | Streaming |
| Multi-device delivery | Every device attached to the conversation receives the same messages in realtime. | Streaming |
| Bidirectional control | Cancel, steer, and interrupt the agent from any client, on the same channel as the response. | Streaming |
| Run tracking | Run lifecycle events say which client started a run and when it ended. view.runs() keeps the tally for you. | Durable session |
| Conversation branching | Edit and regenerate create forks in the conversation tree rather than destructive replacements. | Durable session |
| Approval gates that reach the user anywhere | Pending tool approvals persist on the session until someone acts on them. | Durable session |
| History and replay | Page the channel backwards on reconnect, page refresh, or a new device joining. | Streaming |
| Token compaction | Reconnecting clients receive accumulated responses rather than a replay of every token. | Streaming |
Where they connect
On the server, AI Transport replaces the HTTP response with run.pipe(). To forward a Responses stream over HTTP yourself, you serialise each event into a Server-Sent Events response by hand:
1
2
3
4
5
6
7
8
9
10
11
12
const encoder = new TextEncoder();
return new Response(
new ReadableStream({
async start(controller) {
for await (const event of stream) {
controller.enqueue(encoder.encode(`data: ${JSON.stringify(event)}\n\n`));
}
controller.close();
},
}),
{ headers: { 'Content-Type': 'text/event-stream' } },
);With AI Transport, a raw ResponseStreamEvent is already valid codec output and run.pipe takes the SDK's async iterable directly, so the same stream reaches a durable session in two lines, and every subscribed client can resume, sync, and cancel it:
1
2
const { reason } = await run.pipe(stream);
await run.end({ reason });The integration has three pieces:
ResponsesSessionCodecfrom@ably/ai-transport/openaiencodes eachResponseStreamEventas Ably messages and reassembles them intoOpenAIMessages on the client. The encoding half is the wire codec,ResponsesCodec, which it is built on. It streams assistant text, refusals, reasoning, and function-call arguments, and synthesises the missing start events for a client that attaches mid-stream. The codec passes the raw events through, so the wire tracks OpenAI's own event model.createAgentSession({ client, channelName, codec: ResponsesSessionCodec })from@ably/ai-transportconstructs the agent session bound to the channel from the invocation.run.pipe()reads the model's event stream, encodes each event, and publishes the resulting Ably messages.run.abortSignalpasses the client's cancellation to the Responses API request.
To feed the next turn, toResponsesInput flattens the conversation read from run.view back into the Responses input array. Each stored OpenAIMessage already holds valid Responses input items, so toResponsesInput passes them straight through:
1
2
3
4
5
6
import { toResponsesInput } from '@ably/ai-transport/openai';
while (run.view.hasOlder()) {
await run.view.loadOlder();
}
const input = toResponsesInput(run.view.getMessages().map(({ message }) => message));Run server-side tools
A Responses stream never carries a function call's output. OpenAI surfaces tool output only as model input on the next turn, so the codec adds an event of its own, function_call_output, for the agent to publish after it runs a tool.
Server-executed tools do not suspend the run. The agent runs an agentic loop: call the Responses API, and if the model emits function calls, run them, append the model's output items and the tool outputs to the input, and call it again. The loop continues until the model produces a reply with no tool calls. Each unit of work publishes under its own run.pipe, so a run that calls one tool produces three messages: the turn that emitted the calls, the tool outputs, and the final text turn.
For reasoning models, the loop must re-append the whole turn's output items, including the reasoning items that preceded a function call, since reasoning models expect that reasoning to travel with the call on the next request. The runnable demo implements the full loop.
Run client-side tools and approvals
The codec also carries the client-driven half of tool calling: a client can execute a tool in the browser and publish the result, report a tool failure, or answer a human approval prompt. Gating a call on a human decision needs the codec's second added event, tool-approval-request, which the Responses API has no equivalent for. The suspend and resume mechanics belong to the transport rather than the codec. The agent calls run.suspend() to wait for a client, the client publishes its input on the same runId, and a continuation resumes the run.
ResponsesSessionCodec exposes the full well-known factory set. You address each client input to the assistant message that holds the function_call, and key each payload by the OpenAI snake_case call_id:
1
2
3
4
5
6
7
8
9
10
import { ResponsesSessionCodec } from '@ably/ai-transport/openai';
// A client-run tool succeeded.
await view.send(ResponsesSessionCodec.createToolResult(codecMessageId, { call_id, output }), { runId });
// A client-run tool failed. The message becomes the output the model sees next turn.
await view.send(ResponsesSessionCodec.createToolResultError(codecMessageId, { call_id, message }), { runId });
// A user approved or denied a gated tool.
await view.send(ResponsesSessionCodec.createToolApprovalResponse(codecMessageId, { call_id, approved, reason }), { runId });The Responses function_call_output item has no field for an approval decision or an error, so the codec holds that render-only state out of band on OpenAIMessage.toolCallStates, a map keyed by call_id. toResponsesInput never reads that state, so it cannot reach the model.
Publish an approval request on the call's own message
To gate a tool on a human decision, the agent publishes the codec's own tool-approval-request output, carrying the call_id, the tool name, and the arguments so your UI can render the prompt from the request alone, before the function_call finishes streaming.
The agent publishes the request at the end of the model turn's own run.pipe, so it carries the same codec-message-id as the function_call it gates. The pending approval state, the client's decision, and the function_call then merge onto one message. Published as a separate message, the request strands its pending state on a message the client's response never amends, so the approval prompt never resolves and the agent never sees the call as approved.
1
2
3
4
5
6
7
8
9
10
11
// Agent: the model turn, then an approval request per gated call, on one pipe.
async function* modelTurnWithGateRequests(input, signal, turn) {
yield* modelTurnStream(input, signal, turn);
for (const call of turn.calls) {
if (needsApproval(call.name)) {
yield { type: 'tool-approval-request', call_id: call.call_id, name: call.name, arguments: call.arguments };
}
}
}
await run.pipe(modelTurnWithGateRequests(input, run.abortSignal, turn));Wait for the run to suspend, then wake the agent
The client waits until the run reports suspended before publishing. One model turn can emit a server tool and a client tool on the same message, and resuming while the run is still active races the run's own output, whose server-tool result has not been merged in yet. The provider then rejects the resumed request for the missing output. The run flips to suspended once the agent pauses it awaiting client input.
The client answers every open call before waking the agent. The model input must carry a matching output for every open function_call, so a turn that emits two gated calls needs both answers before either wakes the agent. The client reads the outstanding calls with unansweredCalls. The client's own resolution is wire-only, since the transport skips the optimistic merge for an input targeting an existing message, so your client must track the call_ids it has answered instead of waiting for them to appear in the view.
Publishing an input resumes nothing on its own. As everywhere else in AI Transport, the client publishes and then POSTs the invocation to wake the agent, in that order: the agent reads the conversation off the channel, so the resolution has to be there before the POST arrives.
1
2
3
4
5
6
7
8
9
10
11
12
13
import { unansweredCalls } from '@ably/ai-transport/openai';
const target = view.runOf(codecMessageId);
if (target?.status !== 'suspended') return;
answered.add(call_id);
const runMessages = view.messages
.filter((entry) => view.runOf(entry.codecMessageId)?.runId === target.runId)
.map((entry) => entry.message);
const answeredTheLastCall = unansweredCalls(runMessages).every((call) => answered.has(call.call_id));
const run = await view.send([input], { runId: target.runId });
if (answeredTheLastCall) await wakeAgent(run);Run an approved call on resume
An approval records the user's decision without producing the tool's output, so when the user approves a gated call the conversation holds a function_call with no function_call_output. The agent runs the tool server-side on resume, before the next model turn. It reads those calls with approvedUnexecutedCalls, publishes each output as its own message, and feeds them back into the input:
1
2
3
4
5
6
7
8
import { approvedUnexecutedCalls } from '@ably/ai-transport/openai';
const approved = approvedUnexecutedCalls(priorMessages);
if (approved.length > 0) {
const { items, events } = runToolCalls(approved);
await run.pipe(outputStream(events));
input.push(...items);
}A denial needs no server execution, because the codec resolves it with a rejection function_call_output on the client. The run is still suspended, so the client must wake it to continue.
The third correlation reader, resolvedCallIds, returns the call_ids that already carry an output, which your UI can use to skip the calls it has already shown an answer for.
Scope and trade-offs
The OpenAI codec is in early preview. It covers the shapes the Responses API streams: assistant text, refusals, reasoning (summary and raw), and server-executed function calls. It also covers the client-driven half of tool calling: client-executed tools, tool failures, and human approvals.
Outside that scope today: hosted tools (web and file search, code interpreter, image generation, MCP, custom tools), audio, and rich prompt content. A user prompt's content parts must all be input_text; an input_image or input_file part throws at the encoder.
The codec's event inventory is total, so an event outside it throws at the encoder instead of being dropped silently. An agent that enables a hosted tool must filter that tool's events, and the output_text annotations those tools cite, out of the stream before piping it.
The codec transmits the raw Responses events rather than a normalised abstraction. The wire therefore tracks OpenAI's own event model, which keeps stored items round-trippable to the Responses API but ties a conversation to the Responses shape. To integrate a different model provider behind one abstraction, use Vercel AI SDK Core instead.
Read next
- Get started with OpenAI: build a working app.
- OpenAI codec reference: both codecs, their methods, tool payloads, and types.
- Conversation helpers reference:
toResponsesInputand the correlation readers. - Codec architecture: how a codec translates a framework's events into Ably messages.
- Tool calling: the tool-calling model across the SDK.
- Agent session API reference: every method and property on the agent session.