Vercel WDK
Run the agent side of a conversation as a Vercel Workflow in the same deployment. The workflow opens the run and runs each model call as its own durable WDK step, a retry supersedes the failed attempt on the session, and the user's stream never breaks.
What you build
A Next.js chat app where:
- Every user turn starts a Vercel Workflow Development Kit (WDK) workflow inside the same Next.js deployment. The workflow executes the agent loop as a sequence of WDK steps, each running as its own process.
- The workflow opens the run in one step, then runs each model call and each tool as its own step. Every model response and tool result is published as an AI Transport step, identified by the WDK step id.
- When a step crashes, WDK retries it as a fresh process under the same step id, so the retry's output supersedes the failed attempt's on the session and the user sees one clean result rather than a duplicate.
- The client experience is identical to the Vercel AI SDK getting-started; only the server-side execution model changes.
What Vercel WDK brings
WDK owns execution. It persists a workflow's progress, retries a failed WDK step, and runs it all inside your Next.js app.
| Feature | Description |
|---|---|
| Workflow durability | The workflow's progress persists as a journal of completed WDK steps. A crash or redeploy resumes the workflow with every finished step's result replayed rather than re-executed. |
| Step retries | A failed WDK step re-runs as a fresh process, up to three retries by default. Tune the count per step with maxRetries, control backoff by throwing RetryableError with a retryAfter, or fail immediately with FatalError. The retry keeps the same WDK step id. |
| In-app runtime | Workflows compile into your Next.js app through the 'use workflow' and 'use step' directives. There is no separate cluster or worker fleet, and next dev runs workflows locally with no configuration. |
| Observability | npx workflow web opens an inspector showing runs, steps, attempts, and errors. |
What AI Transport adds
AI Transport owns conversation state: what appears in the conversation, how a retry reconciles with the attempt it replaces, and how every connected device observes it.
| Feature | Description |
|---|---|
| Durable sessions | Tokens flow through a session that outlives any single connection. A client reconnects and resumes from where it left off. |
| Multi-device sync | Every device subscribed to the session sees the same conversation in realtime. |
| Bidirectional control | Cancel, steer, and interrupt the agent from any client, over the same session that carries the response. |
| Retry supersede | When WDK retries a step, AgentRun.createStep({ stepId }) causes the retry's output to supersede the failed attempt's on the session rather than append beside it. |
| Cancel routing on the session | A cancel published to the session reaches every step directly, so a stop button in the browser aborts the LLM call inside the in-flight WDK step. |
| History and replay | Load the full conversation on reconnect, page refresh, or new device join. |
A WDK step that produces agent output publishes it as a single AI Transport step, using the WDK step id as that step's stepId. The WDK step that retries and the AI Transport step that supersedes share one id, and getStepMetadata() reads that id inside the WDK step. WDK step ids are stable across retries of the same step and unique across workflow runs, so you can pass them straight into createStep.
Prerequisites
- Node.js 22 or later.
- An Ably account with an API key.
- An Anthropic API key, or any other model provider supported by Vercel AI SDK.
The workflow runtime ships in the workflow package and runs inside next dev.
Install dependencies
Install the AI Transport SDK, the Vercel AI SDK, the workflow runtime, and Next.js:
npm install @ably/ai-transport ably ai @ai-sdk/react @ai-sdk/anthropic workflow zod \
next react react-domSet up authentication
Create an auth endpoint at /api/auth/token that returns an Ably JWT to the client. The endpoint validates the user and signs a token with their client ID and the channel capabilities they need, as described in Set up authentication.
The client below uses authUrl: '/api/auth/token' to fetch tokens from this endpoint.
Configure the channel rule
AI Transport streams each response by appending tokens to a single channel message. That requires the Message annotations, updates, deletes, and appends channel rule (mutableMessages) on the namespace your conversations live on.
In your Ably dashboard, enable Message annotations, updates, deletes, and appends on the conversations namespace, using the dashboard, Control API, or CLI.
Configure Next.js for workflows
Wrap your Next.js config with withWorkflow. It installs the bundler transforms for the 'use workflow' and 'use step' directives and serves the workflow runtime's handler endpoints under /.well-known/workflow/v1/*.
1
2
3
4
5
6
import type { NextConfig } from 'next';
import { withWorkflow } from 'workflow/next';
const nextConfig: NextConfig = {};
export default withWorkflow(nextConfig);In local development the runtime stores run state in .workflow-data/; add that directory to .gitignore.
Build the agent
The agent side is three files under app/workflows/: the workflow that orchestrates the turn, the server tool it calls, and the steps that publish to the session. Create them next.
Workflow
On the agent, create app/workflows/turn.ts. The workflow is the deterministic orchestrator: it holds the run's identity and does no I/O on the session. openRun opens the run; runInference runs each model call, the first and every follow-up. While an inference reports fresh server-tool calls, the workflow dispatches one runTool per call, then loops a follow-up runInference. Every terminal outcome has already been published on the wire by the step that produced it; the workflow only decides whether to schedule more steps.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
import { getWorkflowMetadata } from 'workflow';
import type { InvocationData } from '@ably/ai-transport';
import { failRun, openRun, runInference, runTool, type TurnIds } from './steps';
export async function chatWorkflow(invocation: InvocationData): Promise<void> {
'use workflow';
const { workflowRunId } = getWorkflowMetadata();
let ids: TurnIds | undefined;
try {
ids = await openRun(invocation, workflowRunId);
let outcome = await runInference(invocation, ids);
while (outcome.kind === 'server-tools') {
for (const toolCall of outcome.serverToolCalls) {
await runTool(invocation, ids, toolCall);
}
outcome = await runInference(invocation, ids);
}
} catch (error) {
// Retries exhausted: if a run was opened, end it in error so observers see the run end.
if (ids) {
const message = error instanceof Error ? error.message : String(error);
try {
await failRun(invocation, ids, message);
} catch {
/* best-effort */
}
}
throw error;
}
}A 'use workflow' function must stay deterministic. WDK re-executes it on every wake-up, replaying each awaited step's recorded result instead of re-running it, so anything nondeterministic, and all I/O on the session, belongs inside the steps. Reading getWorkflowMetadata() here is safe because workflowRunId is stable across replays.
Tool
Create app/workflows/tools.ts with one server tool. It generates a whole-dollar price and throws when the price is odd, about half the time, so you can watch WDK retry the tool step in the inspector and see the retry's re-rolled output supersede the failed attempt on the session.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
import { z } from 'zod';
import type { Tool } from 'ai';
export const tools: Record<string, Tool> = {
getStockPrice: {
description: 'Get the current stock price for a ticker symbol.',
inputSchema: z.object({
symbol: z.string().describe('The ticker symbol, for example "AAPL"'),
}),
execute: async ({ symbol }: { symbol: string }) => {
// Intentionally flaky: throws on an odd price so you can watch the retry re-roll it.
const priceUSD = Math.round(50 + Math.random() * 500);
if (priceUSD % 2 !== 0) {
throw new Error(`stock price service returned an odd price (${priceUSD}), retry me`);
}
return { symbol, priceUSD };
},
},
};Steps
Create app/workflows/steps.ts. Each 'use step' function runs as a separate process, a fresh invocation with no shared memory, so every step constructs its own Ably client and AgentSession and reconstructs run state from the session. Only serializable values cross the workflow-to-step boundary; the client and session never do.
Three structural rules keep the cross-process lifecycle sound:
- The step that produces an outcome publishes the matching run lifecycle event in the same session it streamed with.
- Failure paths detach rather than end, leaving the run open on the wire so a WDK retry can adopt it and publish a superseding attempt under the same
stepId. - Open the step before you read the view. A retry re-enters the step with the same
stepId, sostep.start()re-emitsai-step-startunder a higher channel serial that supersedes the failed attempt, and it merges that superseding start into the read model behindrun.viewbefore it resolves. Readrun.viewfirst and the half-streamed response the retry is replacing sits at the end of the prompt as an assistant prefill, which real providers reject.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
import * as Ably from 'ably';
import { convertToModelMessages, stepCountIs, streamText } from 'ai';
import { getStepMetadata } from 'workflow';
import { anthropic } from '@ai-sdk/anthropic';
import { ErrorCode, Invocation, type InvocationData } from '@ably/ai-transport';
import {
createAgentSession,
pendingToolCalls,
stripToolExecutes,
vercelRunOutcome,
} from '@ably/ai-transport/vercel';
import { tools } from './tools';
export interface TurnIds {
runId: string;
invocationId: string;
}
export interface ToolCallInfo {
toolCallId: string;
toolName: string;
input: unknown;
}
// server-tools is the only non-terminal kind; the other kinds are already published.
export type InferenceOutcome =
| { kind: 'complete' }
| { kind: 'suspend' }
| { kind: 'cancelled' }
| { kind: 'error'; errorMessage: string }
| { kind: 'server-tools'; serverToolCalls: ToolCallInfo[] };
type AgentSession = ReturnType<typeof createAgentSession>;
type AgentRun = ReturnType<AgentSession['adoptRun']> | ReturnType<AgentSession['createRun']>;
// A fresh session per step. detach (not end) leaves any open run on the channel
// for the next step, or for a retry to adopt.
async function withAgentSession<T>(
channelName: string,
body: (session: AgentSession) => Promise<T>,
): Promise<T> {
const client = new Ably.Realtime({ key: process.env.ABLY_API_KEY! });
const session = createAgentSession({ client, channelName });
try {
await session.connect();
return await body(session);
} finally {
try {
await session.detach();
} catch {
/* best-effort */
}
client.close();
}
}
// Open the run and return its ids. No model call here; the first inference is its own step.
export async function openRun(invocationData: InvocationData, workflowRunId: string): Promise<TurnIds> {
'use step';
const invocation = Invocation.fromJSON(invocationData);
const invocationId = `inv:${workflowRunId}`;
return withAgentSession(invocation.sessionName, async (session) => {
// Pin the ids to the replay-stable workflow run id so a retry re-enters the same run.
const run = session.createRun(invocation, { runId: `run:${workflowRunId}`, invocationId });
// Drain history so the trigger folds in, then open (or resume) the run.
while (run.view.hasOlder()) await run.view.loadOlder();
await run.start();
return { runId: run.runId, invocationId };
});
}
// Adopt the open run, run one model call as a step, and publish its terminal.
export async function runInference(invocationData: InvocationData, ids: TurnIds): Promise<InferenceOutcome> {
'use step';
const invocation = Invocation.fromJSON(invocationData);
return withAgentSession(invocation.sessionName, async (session) => {
const { stepId } = getStepMetadata();
const run = session.adoptRun(invocation, ids);
await run.load({ timeoutMs: 15_000 });
while (run.view.hasOlder()) await run.view.loadOlder();
const step = run.createStep({ stepId });
await step.start();
const conversation = run.view.getMessages().map((entry) => entry.message);
const result = streamText({
model: anthropic('claude-sonnet-4-20250514'),
messages: await convertToModelMessages(conversation),
tools: stripToolExecutes(tools),
abortSignal: run.abortSignal,
// The workflow drives the loop; this call runs one step only.
stopWhen: stepCountIs(1),
});
const pipeResult = await step.pipe(result.toUIMessageStream());
const outcome = await vercelRunOutcome(pipeResult, result.finishReason);
await step.end(outcome.reason === 'error' ? { reason: 'failed' } : undefined);
let inference: InferenceOutcome;
if (outcome.reason === 'error') {
inference = { kind: 'error', errorMessage: outcome.error.message };
} else if (outcome.reason === 'cancelled') {
inference = { kind: 'cancelled' };
} else if (outcome.reason === 'complete') {
inference = { kind: 'complete' };
} else {
// Server tools (execute in the registry) run as steps; anything else suspends.
const serverToolCalls = pendingToolCalls(run.messages)
.filter((call) => typeof tools[call.toolName]?.execute === 'function')
.map((call) => ({ toolCallId: call.toolCallId, toolName: call.toolName, input: call.input }));
inference = serverToolCalls.length > 0 ? { kind: 'server-tools', serverToolCalls } : { kind: 'suspend' };
}
await publishTerminal(run, inference);
return inference;
});
}
// Execute one server tool and publish its result as a step. A throw retries under the same stepId.
export async function runTool(invocationData: InvocationData, ids: TurnIds, toolCall: ToolCallInfo): Promise<void> {
'use step';
const invocation = Invocation.fromJSON(invocationData);
await withAgentSession(invocation.sessionName, async (session) => {
const run = session.adoptRun(invocation, ids);
await run.load({ timeoutMs: 15_000 });
const step = run.createStep({ stepId: getStepMetadata().stepId });
await step.start();
// Look the tool up by the model-provided name and narrow it to the callable shape.
const tool = tools[toolCall.toolName] as { execute?: (input: unknown) => Promise<unknown> };
if (!tool?.execute) throw new Error(`tool '${toolCall.toolName}' has no execute`);
const output = await tool.execute(toolCall.input);
await step.send({ type: 'tool-output-available', toolCallId: toolCall.toolCallId, output });
await step.end();
});
}
// The workflow's failure terminal, run from its catch when a step exhausts retries.
// A load() rejection means the run is already gone; nothing to clean up.
export async function failRun(invocationData: InvocationData, ids: TurnIds, errorMessage: string): Promise<void> {
'use step';
const invocation = Invocation.fromJSON(invocationData);
await withAgentSession(invocation.sessionName, async (session) => {
const run = session.adoptRun(invocation, ids);
try {
await run.load({ timeoutMs: 15_000 });
} catch {
return;
}
await run.end({ reason: 'error', error: new Ably.ErrorInfo(errorMessage, ErrorCode.RunResponseStreamFailed, 500) });
});
}
// Cleanup is best-effort and must not cascade: one attempt, no retries.
failRun.maxRetries = 0;
// Publish the lifecycle event the outcome implies; server-tools publishes nothing.
async function publishTerminal(run: AgentRun, outcome: InferenceOutcome): Promise<void> {
switch (outcome.kind) {
case 'server-tools':
return;
case 'suspend':
await run.suspend();
return;
case 'error':
await run.end({ reason: 'error', error: new Ably.ErrorInfo(outcome.errorMessage, ErrorCode.RunResponseStreamFailed, 500) });
return;
default:
await run.end({ reason: outcome.kind });
}
}openRun opens the run and returns its ids without calling the model. Splitting it from the first inference keeps each step independently retryable: an inference retry never re-opens the run, and the ids reach the workflow before any inference, so failRun can still end the run if an inference later exhausts its retries.
runInference adopts the run and runs one model call. stripToolExecutes and stopWhen: stepCountIs(1) keep the Vercel AI SDK from running its own tool loop, so the model emits its tool-call parts and stops, and the workflow executes the loop instead.
Route the tool calls
pendingToolCalls reports the fresh calls the model just emitted, the parts in input-available state. It does not classify them; your registry does, by whether the named tool has an execute:
| Tool kind | How the model call sees it | How the workflow resolves it |
|---|---|---|
| Server tool | An execute in your registry, stripped for the model call | The workflow schedules one runTool step per call. The step runs execute, publishes the result as a tool-output-available message under its own step, then a follow-up inference feeds it back to the model. |
| Client tool | No execute; the browser owns the work | The inference finds no server call to dispatch, so the run suspends. The client runs the tool and posts the result, which starts a continuation. |
| Approval-gated tool | A needsApproval predicate that stripToolExecutes preserves, with its execute stripped like any server tool | The model emits a tool-approval-request part and the run suspends for the user's decision. Once approved, the call runs as a server tool in its own tool step. |
A runTool step throws if its tool has no execute, so the classification and the dispatch stay in agreement.
Approvals resolve on a separate helper. A continuation triggered by a tool-approval-response dispatches the approved call directly with approvedPendingToolCalls rather than replaying the approval through the model, because feeding an approved-then-answered tool pair back through streamText is unreliable on real providers. approvedPendingToolCalls excludes a denied approval, so a denied call never reaches the tool path. Tool calling covers the client-executed and approval-gated flows in full.
Each inference ends with vercelRunOutcome, which maps the stream's finishReason to a terminal or to a pause when the model stopped on tool calls. The step publishes ai-run-end or ai-run-suspend inline for a terminal outcome, and for a server-tool outcome it publishes nothing and leaves the run open for the next tool step to adopt.
Create the agent route
On the agent, create app/api/chat/route.ts. The route starts a workflow and returns immediately; the reply reaches the client over the session, so the response body is informational only.
1
2
3
4
5
6
7
8
9
import { start } from 'workflow/api';
import type { InvocationData } from '@ably/ai-transport';
import { chatWorkflow } from '../../workflows/turn';
export async function POST(req: Request): Promise<Response> {
const invocation = (await req.json()) as InvocationData;
const run = await start(chatWorkflow, [invocation]);
return Response.json({ workflowRunId: run.runId });
}run.runId here is the WDK workflow run id rather than the AI Transport runId; openRun mints that on the session. A continuation POST (a tool result, a regenerate, a suspend resume) starts a fresh workflow the same way, and its openRun step resumes the existing run.
Create the chat component
The client is identical to the Vercel AI SDK getting-started. The ChatTransport POSTs to /api/chat; whether the server side is a single streamText call or a WDK workflow is invisible to the client.
Wire it together
The page wrapper is identical to the Vercel AI SDK getting-started. Use the same Providers and ChatTransportProvider setup.
Run the app
One process runs everything. Next.js serves the app, and the workflow runtime executes workflows and steps inside it:
npm run devTo watch workflows execute, open the inspector in a second terminal:
npx workflow webOpen the app at http://localhost:3000. Every user turn appears as a new workflow run in the inspector; each step's output is a step on the session. Ask for a stock price a few times: about half the tool steps fail and retry, re-rolling the price, and the conversation settles on the retried output with no duplicate message.
What happens when you send a message
sendMessage({ text })publishes the user input on the session and POSTs an invocation to/api/chat, which starts a WDK workflow.openRuncallssession.createRunwith the run id pinned to the workflow run id, drains history so the trigger merges in, and publishesai-run-startwithrun.start(). It returns the run's ids and runs no model. The workflow then callsrunInference, which adopts the run withsession.adoptRun, opens a step undergetStepMetadata().stepId, streams the model response, and publishes the outcome's terminal inline. When the model asks for a server tool, the workflow schedulesrunToolper call (each adopts the run and publishes the result as its own step), then loops back intorunInferencefor the follow-up, which publishesai-run-end.- If a step throws, WDK re-runs it as a fresh process under the same step id.
getStepMetadata().stepIdis stable across retries, so the retry'sai-step-startsupersedes the failed attempt's output on the session. The user sees only the retried step's output. ThegetStockPricetool throws on odd prices to make this visible on the tool step. - Cancel messages arrive on the session rather than through a workflow API. The in-flight step's own session routes them to
run.abortSignal, which flows into the LLM call.
This guide's happy path publishes each run terminal inside the step that produced it. When a step exhausts its retries instead, the workflow's catch schedules failRun to close the run once with run.end({ reason: 'error' }) so every observer sees the run end; maxRetries = 0 keeps that final write to a single attempt. The catch holds the run's ids only once openRun has returned them, and openRun publishes ai-run-start before it returns. A retry re-enters the same run, because the run id is pinned to the workflow run id. If every attempt fails after that publish has landed, the run is open on the channel and the catch has no ids to end it.
Route cancel messages through the session
A cancel does not travel through a workflow-level signal. Mid-flight control between a client and a running step is the transport's job, so a clientRun.cancel() in the browser publishes ai-cancel on the Ably channel. The in-flight WDK step's own AgentSession subscribes to that channel and routes the cancel to run.abortSignal:
Because each WDK step constructs its own session and routes its own cancel messages, run.abortSignal fires in whichever step is subscribed when the cancel lands. An inference step passes that signal into its model call, aborts, and publishes the cancelled terminal inline, and the workflow stops scheduling further steps. A cancel that lands while a tool step is in flight fires that step's run.abortSignal too, so pass that signal into any long-running tool work.
Suspend and resume across workflows
The step holding the run calls run.suspend(), publishes ai-run-suspend, and returns, and the workflow completes. When the client posts a continuation invocation, whether a tool result or an approval response, the agent route starts a fresh workflow. That workflow's openRun step calls createRun as usual, and the SDK reads the existing runId from the continuation's trigger and publishes ai-run-resume rather than ai-run-start.
The invocation's HTTP body carries no conversation content, only the id of the event the client just published on the session, and the agent reads that event from the session to learn what to do. Each HTTP POST starts one workflow run, so continuations are new workflow runs on the same AI Transport runId rather than resumes of the original workflow.
WDK's own hooks can hold a workflow open while it waits for external input. This design ends the workflow instead, because the continuation arrives as a fresh POST whose trigger already points at the event on the session that resumes the run. The client stays identical either way.
Scope and trade-offs
Vercel WDK is intentionally focused on execution durability. By design it does not model conversation state or route signals between clients and a running step. AI Transport adds both without changing how you write workflows. The workflow decides what runs and when, and AI Transport decides what appears in the conversation and to whom.
Boundaries to keep in mind:
- The workflow function is a deterministic orchestrator. WDK re-executes it on every wake-up, replaying recorded step results, and rejects I/O inside it. All work on the session happens inside steps.
- Only serializable values cross the workflow-to-step boundary. An
Ably.Realtimeclient or anAgentSessioncannot be passed between steps, so each step constructs its own and reconstructs run state from the session withadoptRunandload. - Durability applies at the step boundary rather than mid-chunk of an LLM stream. WDK re-runs the whole step, its step opens fresh under the same
stepId, and the retry's output supersedes the failed attempt. - The catch holds the run's ids only once
openRunhas returned them, so a turn whoseopenRunfails every attempt strands no run, because nothing was opened. Alert on persistentopenRunfailure so a turn that never starts does not pass silently.
Explore next
- Durable execution: the same pattern, generalised to any workflow engine.
- Temporal: the same job on a dedicated workflow engine.
- Runs and steps: the unit a WDK step publishes.
- AgentSession API reference:
adoptRun,createStep,detach, andend. - Tool calling: server-executed, client-executed, and approval-gated tools.