Vercel WDK

Run the agent side of a conversation as a Vercel Workflow in the same deployment. The workflow opens the run and runs each model call as its own durable WDK step, a retry supersedes the failed attempt on the session, and the user's stream never breaks.

What you build

A Next.js chat app where:

  • Every user turn starts a Vercel Workflow Development Kit (WDK) workflow inside the same Next.js deployment. The workflow executes the agent loop as a sequence of WDK steps, each running as its own process.
  • The workflow opens the run in one step, then runs each model call and each tool as its own step. Every model response and tool result is published as an AI Transport step, identified by the WDK step id.
  • When a step crashes, WDK retries it as a fresh process under the same step id, so the retry's output supersedes the failed attempt's on the session and the user sees one clean result rather than a duplicate.
  • The client experience is identical to the Vercel AI SDK getting-started; only the server-side execution model changes.
Diagram of a Vercel WDK workflow executing one AI Transport run on the Ably channel. Opening the run publishes ai-run-start with no step; a separate step for the first model call, each tool call, and each follow-up publishes its output as an AI Transport step, A then B then C, using the WDK step id as the step id. One step's attempt crashes and its retry re-runs under the same stepId, superseding the failed attempt, so the run ends with a single clean set of steps that every client sees.

What Vercel WDK brings

WDK owns execution. It persists a workflow's progress, retries a failed WDK step, and runs it all inside your Next.js app.

FeatureDescription
Workflow durabilityThe workflow's progress persists as a journal of completed WDK steps. A crash or redeploy resumes the workflow with every finished step's result replayed rather than re-executed.
Step retriesA failed WDK step re-runs as a fresh process, up to three retries by default. Tune the count per step with maxRetries, control backoff by throwing RetryableError with a retryAfter, or fail immediately with FatalError. The retry keeps the same WDK step id.
In-app runtimeWorkflows compile into your Next.js app through the 'use workflow' and 'use step' directives. There is no separate cluster or worker fleet, and next dev runs workflows locally with no configuration.
Observabilitynpx workflow web opens an inspector showing runs, steps, attempts, and errors.

What AI Transport adds

AI Transport owns conversation state: what appears in the conversation, how a retry reconciles with the attempt it replaces, and how every connected device observes it.

FeatureDescription
Durable sessionsTokens flow through a session that outlives any single connection. A client reconnects and resumes from where it left off.
Multi-device syncEvery device subscribed to the session sees the same conversation in realtime.
Bidirectional controlCancel, steer, and interrupt the agent from any client, over the same session that carries the response.
Retry supersedeWhen WDK retries a step, AgentRun.createStep({ stepId }) causes the retry's output to supersede the failed attempt's on the session rather than append beside it.
Cancel routing on the sessionA cancel published to the session reaches every step directly, so a stop button in the browser aborts the LLM call inside the in-flight WDK step.
History and replayLoad the full conversation on reconnect, page refresh, or new device join.

A WDK step that produces agent output publishes it as a single AI Transport step, using the WDK step id as that step's stepId. The WDK step that retries and the AI Transport step that supersedes share one id, and getStepMetadata() reads that id inside the WDK step. WDK step ids are stable across retries of the same step and unique across workflow runs, so you can pass them straight into createStep.

Prerequisites

  • Node.js 22 or later.
  • An Ably account with an API key.
  • An Anthropic API key, or any other model provider supported by Vercel AI SDK.

The workflow runtime ships in the workflow package and runs inside next dev.

Install dependencies

Install the AI Transport SDK, the Vercel AI SDK, the workflow runtime, and Next.js:

npm install @ably/ai-transport ably ai @ai-sdk/react @ai-sdk/anthropic workflow zod \
  next react react-dom

Set up authentication

Create an auth endpoint at /api/auth/token that returns an Ably JWT to the client. The endpoint validates the user and signs a token with their client ID and the channel capabilities they need, as described in Set up authentication.

The client below uses authUrl: '/api/auth/token' to fetch tokens from this endpoint.

Configure the channel rule

AI Transport streams each response by appending tokens to a single channel message. That requires the Message annotations, updates, deletes, and appends channel rule (mutableMessages) on the namespace your conversations live on.

In your Ably dashboard, enable Message annotations, updates, deletes, and appends on the conversations namespace, using the dashboard, Control API, or CLI.

Configure Next.js for workflows

Wrap your Next.js config with withWorkflow. It installs the bundler transforms for the 'use workflow' and 'use step' directives and serves the workflow runtime's handler endpoints under /.well-known/workflow/v1/*.

JavaScript

1

2

3

4

5

6

import type { NextConfig } from 'next';
import { withWorkflow } from 'workflow/next';

const nextConfig: NextConfig = {};

export default withWorkflow(nextConfig);

In local development the runtime stores run state in .workflow-data/; add that directory to .gitignore.

Build the agent

The agent side is three files under app/workflows/: the workflow that orchestrates the turn, the server tool it calls, and the steps that publish to the session. Create them next.

Workflow

On the agent, create app/workflows/turn.ts. The workflow is the deterministic orchestrator: it holds the run's identity and does no I/O on the session. openRun opens the run; runInference runs each model call, the first and every follow-up. While an inference reports fresh server-tool calls, the workflow dispatches one runTool per call, then loops a follow-up runInference. Every terminal outcome has already been published on the wire by the step that produced it; the workflow only decides whether to schedule more steps.

JavaScript

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

26

27

28

29

30

31

32

import { getWorkflowMetadata } from 'workflow';
import type { InvocationData } from '@ably/ai-transport';
import { failRun, openRun, runInference, runTool, type TurnIds } from './steps';

export async function chatWorkflow(invocation: InvocationData): Promise<void> {
  'use workflow';
  const { workflowRunId } = getWorkflowMetadata();

  let ids: TurnIds | undefined;
  try {
    ids = await openRun(invocation, workflowRunId);
    let outcome = await runInference(invocation, ids);

    while (outcome.kind === 'server-tools') {
      for (const toolCall of outcome.serverToolCalls) {
        await runTool(invocation, ids, toolCall);
      }
      outcome = await runInference(invocation, ids);
    }
  } catch (error) {
    // Retries exhausted: if a run was opened, end it in error so observers see the run end.
    if (ids) {
      const message = error instanceof Error ? error.message : String(error);
      try {
        await failRun(invocation, ids, message);
      } catch {
        /* best-effort */
      }
    }
    throw error;
  }
}

A 'use workflow' function must stay deterministic. WDK re-executes it on every wake-up, replaying each awaited step's recorded result instead of re-running it, so anything nondeterministic, and all I/O on the session, belongs inside the steps. Reading getWorkflowMetadata() here is safe because workflowRunId is stable across replays.

Tool

Create app/workflows/tools.ts with one server tool. It generates a whole-dollar price and throws when the price is odd, about half the time, so you can watch WDK retry the tool step in the inspector and see the retry's re-rolled output supersede the failed attempt on the session.

JavaScript

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

import { z } from 'zod';
import type { Tool } from 'ai';

export const tools: Record<string, Tool> = {
  getStockPrice: {
    description: 'Get the current stock price for a ticker symbol.',
    inputSchema: z.object({
      symbol: z.string().describe('The ticker symbol, for example "AAPL"'),
    }),
    execute: async ({ symbol }: { symbol: string }) => {
      // Intentionally flaky: throws on an odd price so you can watch the retry re-roll it.
      const priceUSD = Math.round(50 + Math.random() * 500);
      if (priceUSD % 2 !== 0) {
        throw new Error(`stock price service returned an odd price (${priceUSD}), retry me`);
      }
      return { symbol, priceUSD };
    },
  },
};

Steps

Create app/workflows/steps.ts. Each 'use step' function runs as a separate process, a fresh invocation with no shared memory, so every step constructs its own Ably client and AgentSession and reconstructs run state from the session. Only serializable values cross the workflow-to-step boundary; the client and session never do.

Three structural rules keep the cross-process lifecycle sound:

  • The step that produces an outcome publishes the matching run lifecycle event in the same session it streamed with.
  • Failure paths detach rather than end, leaving the run open on the wire so a WDK retry can adopt it and publish a superseding attempt under the same stepId.
  • Open the step before you read the view. A retry re-enters the step with the same stepId, so step.start() re-emits ai-step-start under a higher channel serial that supersedes the failed attempt, and it merges that superseding start into the read model behind run.view before it resolves. Read run.view first and the half-streamed response the retry is replacing sits at the end of the prompt as an assistant prefill, which real providers reject.
JavaScript

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

26

27

28

29

30

31

32

33

34

35

36

37

38

39

40

41

42

43

44

45

46

47

48

49

50

51

52

53

54

55

56

57

58

59

60

61

62

63

64

65

66

67

68

69

70

71

72

73

74

75

76

77

78

79

80

81

82

83

84

85

86

87

88

89

90

91

92

93

94

95

96

97

98

99

100

101

102

103

104

105

106

107

108

109

110

111

112

113

114

115

116

117

118

119

120

121

122

123

124

125

126

127

128

129

130

131

132

133

134

135

136

137

138

139

140

141

142

143

144

145

146

147

148

149

150

151

152

153

154

155

156

157

158

159

160

161

162

163

164

165

166

167

168

169

170

171

172

173

import * as Ably from 'ably';
import { convertToModelMessages, stepCountIs, streamText } from 'ai';
import { getStepMetadata } from 'workflow';
import { anthropic } from '@ai-sdk/anthropic';
import { ErrorCode, Invocation, type InvocationData } from '@ably/ai-transport';
import {
  createAgentSession,
  pendingToolCalls,
  stripToolExecutes,
  vercelRunOutcome,
} from '@ably/ai-transport/vercel';
import { tools } from './tools';

export interface TurnIds {
  runId: string;
  invocationId: string;
}

export interface ToolCallInfo {
  toolCallId: string;
  toolName: string;
  input: unknown;
}

// server-tools is the only non-terminal kind; the other kinds are already published.
export type InferenceOutcome =
  | { kind: 'complete' }
  | { kind: 'suspend' }
  | { kind: 'cancelled' }
  | { kind: 'error'; errorMessage: string }
  | { kind: 'server-tools'; serverToolCalls: ToolCallInfo[] };

type AgentSession = ReturnType<typeof createAgentSession>;
type AgentRun = ReturnType<AgentSession['adoptRun']> | ReturnType<AgentSession['createRun']>;

// A fresh session per step. detach (not end) leaves any open run on the channel
// for the next step, or for a retry to adopt.
async function withAgentSession<T>(
  channelName: string,
  body: (session: AgentSession) => Promise<T>,
): Promise<T> {
  const client = new Ably.Realtime({ key: process.env.ABLY_API_KEY! });
  const session = createAgentSession({ client, channelName });
  try {
    await session.connect();
    return await body(session);
  } finally {
    try {
      await session.detach();
    } catch {
      /* best-effort */
    }
    client.close();
  }
}

// Open the run and return its ids. No model call here; the first inference is its own step.
export async function openRun(invocationData: InvocationData, workflowRunId: string): Promise<TurnIds> {
  'use step';
  const invocation = Invocation.fromJSON(invocationData);
  const invocationId = `inv:${workflowRunId}`;
  return withAgentSession(invocation.sessionName, async (session) => {
    // Pin the ids to the replay-stable workflow run id so a retry re-enters the same run.
    const run = session.createRun(invocation, { runId: `run:${workflowRunId}`, invocationId });

    // Drain history so the trigger folds in, then open (or resume) the run.
    while (run.view.hasOlder()) await run.view.loadOlder();
    await run.start();

    return { runId: run.runId, invocationId };
  });
}

// Adopt the open run, run one model call as a step, and publish its terminal.
export async function runInference(invocationData: InvocationData, ids: TurnIds): Promise<InferenceOutcome> {
  'use step';
  const invocation = Invocation.fromJSON(invocationData);
  return withAgentSession(invocation.sessionName, async (session) => {
    const { stepId } = getStepMetadata();
    const run = session.adoptRun(invocation, ids);
    await run.load({ timeoutMs: 15_000 });
    while (run.view.hasOlder()) await run.view.loadOlder();

    const step = run.createStep({ stepId });
    await step.start();

    const conversation = run.view.getMessages().map((entry) => entry.message);
    const result = streamText({
      model: anthropic('claude-sonnet-4-20250514'),
      messages: await convertToModelMessages(conversation),
      tools: stripToolExecutes(tools),
      abortSignal: run.abortSignal,
      // The workflow drives the loop; this call runs one step only.
      stopWhen: stepCountIs(1),
    });

    const pipeResult = await step.pipe(result.toUIMessageStream());
    const outcome = await vercelRunOutcome(pipeResult, result.finishReason);
    await step.end(outcome.reason === 'error' ? { reason: 'failed' } : undefined);

    let inference: InferenceOutcome;
    if (outcome.reason === 'error') {
      inference = { kind: 'error', errorMessage: outcome.error.message };
    } else if (outcome.reason === 'cancelled') {
      inference = { kind: 'cancelled' };
    } else if (outcome.reason === 'complete') {
      inference = { kind: 'complete' };
    } else {
      // Server tools (execute in the registry) run as steps; anything else suspends.
      const serverToolCalls = pendingToolCalls(run.messages)
        .filter((call) => typeof tools[call.toolName]?.execute === 'function')
        .map((call) => ({ toolCallId: call.toolCallId, toolName: call.toolName, input: call.input }));
      inference = serverToolCalls.length > 0 ? { kind: 'server-tools', serverToolCalls } : { kind: 'suspend' };
    }

    await publishTerminal(run, inference);
    return inference;
  });
}

// Execute one server tool and publish its result as a step. A throw retries under the same stepId.
export async function runTool(invocationData: InvocationData, ids: TurnIds, toolCall: ToolCallInfo): Promise<void> {
  'use step';
  const invocation = Invocation.fromJSON(invocationData);
  await withAgentSession(invocation.sessionName, async (session) => {
    const run = session.adoptRun(invocation, ids);
    await run.load({ timeoutMs: 15_000 });

    const step = run.createStep({ stepId: getStepMetadata().stepId });
    await step.start();
    // Look the tool up by the model-provided name and narrow it to the callable shape.
    const tool = tools[toolCall.toolName] as { execute?: (input: unknown) => Promise<unknown> };
    if (!tool?.execute) throw new Error(`tool '${toolCall.toolName}' has no execute`);
    const output = await tool.execute(toolCall.input);
    await step.send({ type: 'tool-output-available', toolCallId: toolCall.toolCallId, output });
    await step.end();
  });
}

// The workflow's failure terminal, run from its catch when a step exhausts retries.
// A load() rejection means the run is already gone; nothing to clean up.
export async function failRun(invocationData: InvocationData, ids: TurnIds, errorMessage: string): Promise<void> {
  'use step';
  const invocation = Invocation.fromJSON(invocationData);
  await withAgentSession(invocation.sessionName, async (session) => {
    const run = session.adoptRun(invocation, ids);
    try {
      await run.load({ timeoutMs: 15_000 });
    } catch {
      return;
    }
    await run.end({ reason: 'error', error: new Ably.ErrorInfo(errorMessage, ErrorCode.RunResponseStreamFailed, 500) });
  });
}

// Cleanup is best-effort and must not cascade: one attempt, no retries.
failRun.maxRetries = 0;

// Publish the lifecycle event the outcome implies; server-tools publishes nothing.
async function publishTerminal(run: AgentRun, outcome: InferenceOutcome): Promise<void> {
  switch (outcome.kind) {
    case 'server-tools':
      return;
    case 'suspend':
      await run.suspend();
      return;
    case 'error':
      await run.end({ reason: 'error', error: new Ably.ErrorInfo(outcome.errorMessage, ErrorCode.RunResponseStreamFailed, 500) });
      return;
    default:
      await run.end({ reason: outcome.kind });
  }
}

openRun opens the run and returns its ids without calling the model. Splitting it from the first inference keeps each step independently retryable: an inference retry never re-opens the run, and the ids reach the workflow before any inference, so failRun can still end the run if an inference later exhausts its retries.

runInference adopts the run and runs one model call. stripToolExecutes and stopWhen: stepCountIs(1) keep the Vercel AI SDK from running its own tool loop, so the model emits its tool-call parts and stops, and the workflow executes the loop instead.

Route the tool calls

pendingToolCalls reports the fresh calls the model just emitted, the parts in input-available state. It does not classify them; your registry does, by whether the named tool has an execute:

Tool kindHow the model call sees itHow the workflow resolves it
Server toolAn execute in your registry, stripped for the model callThe workflow schedules one runTool step per call. The step runs execute, publishes the result as a tool-output-available message under its own step, then a follow-up inference feeds it back to the model.
Client toolNo execute; the browser owns the workThe inference finds no server call to dispatch, so the run suspends. The client runs the tool and posts the result, which starts a continuation.
Approval-gated toolA needsApproval predicate that stripToolExecutes preserves, with its execute stripped like any server toolThe model emits a tool-approval-request part and the run suspends for the user's decision. Once approved, the call runs as a server tool in its own tool step.

A runTool step throws if its tool has no execute, so the classification and the dispatch stay in agreement.

Approvals resolve on a separate helper. A continuation triggered by a tool-approval-response dispatches the approved call directly with approvedPendingToolCalls rather than replaying the approval through the model, because feeding an approved-then-answered tool pair back through streamText is unreliable on real providers. approvedPendingToolCalls excludes a denied approval, so a denied call never reaches the tool path. Tool calling covers the client-executed and approval-gated flows in full.

Each inference ends with vercelRunOutcome, which maps the stream's finishReason to a terminal or to a pause when the model stopped on tool calls. The step publishes ai-run-end or ai-run-suspend inline for a terminal outcome, and for a server-tool outcome it publishes nothing and leaves the run open for the next tool step to adopt.

Create the agent route

On the agent, create app/api/chat/route.ts. The route starts a workflow and returns immediately; the reply reaches the client over the session, so the response body is informational only.

JavaScript

1

2

3

4

5

6

7

8

9

import { start } from 'workflow/api';
import type { InvocationData } from '@ably/ai-transport';
import { chatWorkflow } from '../../workflows/turn';

export async function POST(req: Request): Promise<Response> {
  const invocation = (await req.json()) as InvocationData;
  const run = await start(chatWorkflow, [invocation]);
  return Response.json({ workflowRunId: run.runId });
}

run.runId here is the WDK workflow run id rather than the AI Transport runId; openRun mints that on the session. A continuation POST (a tool result, a regenerate, a suspend resume) starts a fresh workflow the same way, and its openRun step resumes the existing run.

Create the chat component

The client is identical to the Vercel AI SDK getting-started. The ChatTransport POSTs to /api/chat; whether the server side is a single streamText call or a WDK workflow is invisible to the client.

Wire it together

The page wrapper is identical to the Vercel AI SDK getting-started. Use the same Providers and ChatTransportProvider setup.

Run the app

One process runs everything. Next.js serves the app, and the workflow runtime executes workflows and steps inside it:

npm run dev

To watch workflows execute, open the inspector in a second terminal:

npx workflow web

Open the app at http://localhost:3000. Every user turn appears as a new workflow run in the inspector; each step's output is a step on the session. Ask for a stock price a few times: about half the tool steps fail and retry, re-rolling the price, and the conversation settles on the retried output with no duplicate message.

What happens when you send a message

  1. sendMessage({ text }) publishes the user input on the session and POSTs an invocation to /api/chat, which starts a WDK workflow.
  2. openRun calls session.createRun with the run id pinned to the workflow run id, drains history so the trigger merges in, and publishes ai-run-start with run.start(). It returns the run's ids and runs no model. The workflow then calls runInference, which adopts the run with session.adoptRun, opens a step under getStepMetadata().stepId, streams the model response, and publishes the outcome's terminal inline. When the model asks for a server tool, the workflow schedules runTool per call (each adopts the run and publishes the result as its own step), then loops back into runInference for the follow-up, which publishes ai-run-end.
  3. If a step throws, WDK re-runs it as a fresh process under the same step id. getStepMetadata().stepId is stable across retries, so the retry's ai-step-start supersedes the failed attempt's output on the session. The user sees only the retried step's output. The getStockPrice tool throws on odd prices to make this visible on the tool step.
  4. Cancel messages arrive on the session rather than through a workflow API. The in-flight step's own session routes them to run.abortSignal, which flows into the LLM call.

This guide's happy path publishes each run terminal inside the step that produced it. When a step exhausts its retries instead, the workflow's catch schedules failRun to close the run once with run.end({ reason: 'error' }) so every observer sees the run end; maxRetries = 0 keeps that final write to a single attempt. The catch holds the run's ids only once openRun has returned them, and openRun publishes ai-run-start before it returns. A retry re-enters the same run, because the run id is pinned to the workflow run id. If every attempt fails after that publish has landed, the run is open on the channel and the catch has no ids to end it.

Route cancel messages through the session

A cancel does not travel through a workflow-level signal. Mid-flight control between a client and a running step is the transport's job, so a clientRun.cancel() in the browser publishes ai-cancel on the Ably channel. The in-flight WDK step's own AgentSession subscribes to that channel and routes the cancel to run.abortSignal:

Diagram of a cancel flowing over the Ably channel. A browser client's cancel publishes ai-cancel on the channel. The in-flight WDK inference step's AgentSession, subscribed to the channel, receives it, matched by run id, and fires run.abortSignal into the model call. The model call aborts, and the step publishes ai-step-end and ai-run-end with reason cancelled.

Because each WDK step constructs its own session and routes its own cancel messages, run.abortSignal fires in whichever step is subscribed when the cancel lands. An inference step passes that signal into its model call, aborts, and publishes the cancelled terminal inline, and the workflow stops scheduling further steps. A cancel that lands while a tool step is in flight fires that step's run.abortSignal too, so pass that signal into any long-running tool work.

Suspend and resume across workflows

The step holding the run calls run.suspend(), publishes ai-run-suspend, and returns, and the workflow completes. When the client posts a continuation invocation, whether a tool result or an approval response, the agent route starts a fresh workflow. That workflow's openRun step calls createRun as usual, and the SDK reads the existing runId from the continuation's trigger and publishes ai-run-resume rather than ai-run-start.

The invocation's HTTP body carries no conversation content, only the id of the event the client just published on the session, and the agent reads that event from the session to learn what to do. Each HTTP POST starts one workflow run, so continuations are new workflow runs on the same AI Transport runId rather than resumes of the original workflow.

WDK's own hooks can hold a workflow open while it waits for external input. This design ends the workflow instead, because the continuation arrives as a fresh POST whose trigger already points at the event on the session that resumes the run. The client stays identical either way.

Scope and trade-offs

Vercel WDK is intentionally focused on execution durability. By design it does not model conversation state or route signals between clients and a running step. AI Transport adds both without changing how you write workflows. The workflow decides what runs and when, and AI Transport decides what appears in the conversation and to whom.

Boundaries to keep in mind:

  • The workflow function is a deterministic orchestrator. WDK re-executes it on every wake-up, replaying recorded step results, and rejects I/O inside it. All work on the session happens inside steps.
  • Only serializable values cross the workflow-to-step boundary. An Ably.Realtime client or an AgentSession cannot be passed between steps, so each step constructs its own and reconstructs run state from the session with adoptRun and load.
  • Durability applies at the step boundary rather than mid-chunk of an LLM stream. WDK re-runs the whole step, its step opens fresh under the same stepId, and the retry's output supersedes the failed attempt.
  • The catch holds the run's ids only once openRun has returned them, so a turn whose openRun fails every attempt strands no run, because nothing was opened. Alert on persistent openRun failure so a turn that never starts does not pass silently.

Explore next