October 1, 2026

Vercel AI SDK tutorial: build a streaming Claude chat in Next.js

Build a streaming chat with structured output and a tool in Next.js using AI SDK 7 and Claude. Server-side key, rate limits, full code.

Recruiting

Tech

Guide

This Vercel AI SDK tutorial builds one working feature end to end: a support chat in Next.js that streams Claude's reply, calls a tool to look up an order, and has a second route that returns a typed triage object. The Anthropic key never leaves the server. It targets AI SDK 7, so none of the v4 to v6 code you'll find in older posts (toDataStreamResponse, handleSubmit, maxSteps) appears here. If you're hiring for this stack instead of building it yourself, start with Hire developers by tech stack: rates, vetting and interview guides.

What you'll build

A support assistant for Northwind Outfitters, a fictional outdoor-gear shop. A customer asks where order A-1042 is, the assistant calls a lookupOrder tool and streams an answer, and a button sends the conversation to a second route that returns {category, urgency, summary}. Plan on roughly 45 minutes. The steps and Full code use TypeScript only.

The transcript and the JSON below are an example of the shape, not recorded output. The order data comes from a fake in-memory table in the tutorial code.

text
You:       Where is order A-1042? It has been 9 days.
Assistant: (calls lookupOrder with { orderId: "A-1042" })
Assistant: A-1042 shipped 9 days ago with NorthPost and is still in transit. Its last
           scan was at a regional sorting hub. I'm sorry about the wait. Want me to
           open a delivery trace?
JSON
{
  "category": "shipping_delay",
  "urgency": "high",
  "summary": "Customer asks about order A-1042, shipped 9 days ago and still in transit."
}

The finished project looks like this:

text
northwind-support/
  app/
    api/
      chat/route.ts       streaming chat route
      triage/route.ts     structured output route
    page.tsx              chat UI
  lib/
    guard.ts              rate limit and input checks
    schema.ts             triage schema
    tools.ts              lookupOrder tool
  .env.local
  .env.example
  package.json

What you need before you start

  • Node.js 22 or later. The AI SDK 7.0 migration guide says AI SDK 7.0 requires it. Next.js itself asks for 20.9 or later per the Next.js installation guide, so the AI SDK's minimum is the one that governs.
  • npm, and an Anthropic API key from the Claude Console.
  • TypeScript and React basics. You don't need to know Next.js route handlers beforehand.
  • A macOS or Linux shell. On Windows, use WSL. The commands use printf, cp and mkdir -p.

The code was built and tested on Next.js 16.3.8 (the version in the Next.js installation guide on 2026-10-01). create-next-app@latest installs whatever is current when you run it. If you want context on what that release changed, Next.js 16.3: what shipped and what Vercel's benchmarks actually show covers it.

Quickstart

Already have a Next.js app and a key? This is the shortest working path. If you're starting from nothing, skip to the steps, which build the whole project from an empty folder. Install the packages:

Bash
npm install ai@7 @ai-sdk/react@4 @ai-sdk/anthropic@4 zod@4

Put ANTHROPIC_API_KEY=your-api-key-here in .env.local, then add app/api/chat/route.ts:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import {
  convertToModelMessages,
  createUIMessageStreamResponse,
  streamText,
  toUIMessageStream,
  type UIMessage,
} from 'ai';

export async function POST(req: Request) {
  const { messages }: { messages: UIMessage[] } = await req.json();

  const result = streamText({
    model: anthropic('claude-sonnet-5-5'),
    messages: await convertToModelMessages(messages),
    maxOutputTokens: 2048,
  });

  return createUIMessageStreamResponse({
    stream: toUIMessageStream({ stream: result.stream }),
  });
}

And replace app/page.tsx:

TypeScript
'use client';

import { useChat } from '@ai-sdk/react';
import { useState } from 'react';

export default function Page() {
  const [input, setInput] = useState('');
  const { messages, sendMessage } = useChat();

  return (
    <main>
      {messages.map((m) => (
        <p key={m.id}>{m.parts.map((p) => (p.type === 'text' ? p.text : ''))}</p>
      ))}
      <form onSubmit={(e) => { e.preventDefault(); sendMessage({ text: input }); setInput(''); }}>
        <input value={input} onChange={(e) => setInput(e.currentTarget.value)} />
      </form>
    </main>
  );
}

Run npm run dev, open http://localhost:3000, type a message and press Enter. Three things to get right before you build on it:

  • Keep the key in a server-side variable. Never prefix it with NEXT_PUBLIC_, because Next.js ships every NEXT_PUBLIC_ variable to the browser.
  • Next.js loads .env.local on its own, and the file is gitignored by default in a new app. Don't paste the key into source files.
  • Always set maxOutputTokens. Without a cap, one chatty answer sets the cost.

How to build a chat feature with the Vercel AI SDK

1. Create the project and install the packages

Create the app with the defaults, install every package the tutorial uses, and set up the env files. Run this in a terminal, in the folder where you keep projects:

Bash
npx create-next-app@latest northwind-support --yes
cd northwind-support
npm install ai@7 @ai-sdk/react@4 @ai-sdk/anthropic@4 zod@4
printf 'ANTHROPIC_API_KEY=\n' > .env.example
cp .env.example .env.local
printf '\n!.env.example\n' >> .gitignore

Every later command runs from northwind-support/. Open .env.local and put your key after the equals sign:

text
ANTHROPIC_API_KEY=your-api-key-here

The --yes flag accepts the installation guide's defaults: TypeScript, Tailwind, ESLint, the App Router, Turbopack, the @/* import alias and an AGENTS.md file. The generated .gitignore ignores .env*, so the last command adds an exception that lets .env.example be committed. .env.local stays ignored.

The packages do separate jobs. ai is the core library with streamText, generateText, tool and the response helpers. @ai-sdk/react holds the useChat hook. @ai-sdk/anthropic is the Claude provider, and the AI SDK's Anthropic provider page documents its default ANTHROPIC_API_KEY variable. zod defines the schemas for tool input and structured output, and AI SDK 7 accepts Zod 3.25.76 or later, or 4.1.8 or later.

A variable without the NEXT_PUBLIC_ prefix is available to route handlers and never gets bundled into client code. The chat page never sees the key. A key in browser code is a published key.

Checkpoint, from northwind-support/:

Bash
npm ls ai @ai-sdk/react @ai-sdk/anthropic zod next react react-dom --depth=0

You should see something like this (an example, your versions will be newer within the same majors):

text
northwind-support@0.1.0
├── @ai-sdk/anthropic@4.0.71
├── @ai-sdk/react@4.0.129
├── ai@7.0.126
├── next@16.3.8
├── react-dom@19.x
├── react@19.x
└── zod@4.6.5

If a package is missing from the list, re-run the npm install line.

2. Stream a Claude reply from a chat route

The route turns the browser's messages into model messages, calls streamText, and returns the result as a UI message stream. Create the folder and the file:

Bash
mkdir -p app/api/chat

Create app/api/chat/route.ts with this complete file. Later steps replace it with a longer version:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import {
  convertToModelMessages,
  createUIMessageStreamResponse,
  streamText,
  toUIMessageStream,
  type UIMessage,
} from 'ai';

const INSTRUCTIONS = [
  'You are the support assistant for Northwind Outfitters, an outdoor-gear shop.',
  'Keep replies under 120 words and say so plainly when you cannot find an order.',
].join(' ');

export async function POST(req: Request) {
  const { messages }: { messages: UIMessage[] } = await req.json();

  let modelMessages;
  try {
    modelMessages = await convertToModelMessages(messages);
  } catch {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  const result = streamText({
    model: anthropic('claude-sonnet-5-5'),
    instructions: INSTRUCTIONS,
    messages: modelMessages,
    maxOutputTokens: 2048,
    abortSignal: req.signal,
  });

  return createUIMessageStreamResponse({
    stream: toUIMessageStream({ stream: result.stream }),
  });
}

This follows the AI SDK's Next.js getting-started guide. convertToModelMessages is async, so it gets awaited, and a malformed message list returns a 400 instead of a 500. The assistant's behaviour goes in instructions, the v7 name for what used to be system. maxOutputTokens caps the cost of one reply. abortSignal: req.signal stops the model call when the browser disconnects, so you don't pay for output nobody reads.

The last line is the v7 change that trips people up. In v6 you called result.toUIMessageStreamResponse(). In v7 the new pair works in two steps: toUIMessageStream converts result.stream to UI messages, and createUIMessageStreamResponse wraps that in an HTTP response. The Reference section has the full rename table.

claude-sonnet-5-5 is the current Sonnet ID in both the provider docs and Anthropic's models overview. Older tutorials use retired IDs, so check the ID first when you copy code. For the Messages API basics underneath, such as what a content block is, How to use the Claude API: build an AI feature step by step covers them without the AI SDK on top.

Checkpoint: start the dev server in one terminal and leave it running.

Bash
npm run dev

In a second terminal, from northwind-support/, send a request in the shape useChat sends:

Bash
curl -N -X POST http://localhost:3000/api/chat \
  -H "content-type: application/json" \
  -d '{"messages":[{"id":"1","role":"user","parts":[{"type":"text","text":"Where is order A-1042? It has been 9 days."}]}]}'

You should see something like this (example output, trimmed, and the deltas differ every run). Each data: line is one chunk of the UI message stream, and the reply arrives as text-delta chunks:

text
data: {"type":"start"}

data: {"type":"start-step"}

data: {"type":"text-start","id":"0"}

data: {"type":"text-delta","id":"0","delta":"I'm sorry about the wait. "}

data: {"type":"text-delta","id":"0","delta":"I can't look up order details yet"}

data: {"type":"text-end","id":"0"}

data: {"type":"finish-step"}

data: {"type":"finish","finishReason":"stop"}

data: [DONE]

If the key is missing or wrong, the request still returns 200, but the stream holds an error chunk instead of text (example output):

text
data: {"type":"start"}

data: {"type":"error","errorText":"An error occurred."}

data: [DONE]

The real cause is in the terminal running npm run dev: Error [AI_APICallError]: API key is invalid. with statusCode: 401. Check .env.local, then stop and restart npm run dev, since Next.js reads the file at startup.

3. Build the chat page with useChat

useChat manages the message list, the request status and the streaming updates, and you supply the input box. In AI SDK 7, useChat doesn't hand you an input value or a handleSubmit. You keep the input in your own useState and call sendMessage({ text: input }).

Replace app/page.tsx with this complete file. The hook POSTs to /api/chat by default, and the transport is passed explicitly here so the endpoint is visible. Its DefaultChatTransport options (api, headers, body, credentials) are in the chatbot UI docs.

TypeScript
'use client';

import { useChat } from '@ai-sdk/react';
import { DefaultChatTransport } from 'ai';
import { useState } from 'react';

const transport = new DefaultChatTransport({ api: '/api/chat' });

export default function Page() {
  const [input, setInput] = useState('Where is order A-1042? It has been 9 days.');
  const { messages, sendMessage, status, error, regenerate, stop } = useChat({ transport });

  const busy = status === 'submitted' || status === 'streaming';

  return (
    <main className="mx-auto max-w-2xl space-y-4 p-6">
      <h1 className="text-xl font-semibold">Northwind Outfitters support</h1>

      {messages.map((m) => (
        <div key={m.id}>
          <strong>{m.role === 'user' ? 'You' : 'Assistant'}</strong>
          {m.parts.map((part, i) => {
            if (part.type === 'text') {
              return (
                <p key={i} className="whitespace-pre-wrap">
                  {part.text}
                </p>
              );
            }
            return null;
          })}
        </div>
      ))}

      {error && (
        <p className="text-red-600">
          {error.message}{' '}
          <button type="button" className="underline" onClick={() => regenerate()}>
            Retry
          </button>
        </p>
      )}

      <form
        className="flex gap-2"
        onSubmit={(e) => {
          e.preventDefault();
          if (busy || input.trim() === '') return;
          sendMessage({ text: input });
          setInput('');
        }}
      >
        <input
          className="flex-1 border p-2"
          value={input}
          onChange={(e) => setInput(e.currentTarget.value)}
          placeholder="Ask about an order"
        />
        <button type="submit" className="border px-3" disabled={busy}>
          Send
        </button>
        <button type="button" className="border px-3" onClick={() => stop()} disabled={!busy}>
          Stop
        </button>
      </form>
    </main>
  );
}

The page renders message.parts, not a single string. A message is a list of parts, and a part's type tells you what it is. Text parts have a text field, and everything else renders as nothing for now. Sonnet 5.5 supports adaptive thinking with a default effort of high (models overview), and some setups return reasoning parts before the text. Code that assumes the first part is text breaks on some requests and works on others, so branch on part.type. The status, stop, regenerate and error values drive the Send, Stop and Retry buttons.

Checkpoint: open http://localhost:3000 (the dev server from step 2 is still running) and press Send on the prefilled question. You should see something like this (an example, since the model's wording varies and it has no order data yet):

text
You
Where is order A-1042? It has been 9 days.
Assistant
I'm sorry about the wait. I can't look up order details yet, so I can't say where A-1042 is.

If a red An error occurred. appears, read the terminal running npm run dev for the real cause. A missing or wrong key in .env.local is the usual one, and shows as AI_APICallError with statusCode: 401.

4. Add the order lookup tool

A tool is a function the model can choose to run. You describe it with tool(), give it a Zod inputSchema, and write an execute function. The model fills in the input, the SDK runs execute, and the result goes back into the conversation. Create the folder and the tool file:

Bash
mkdir -p lib

Create lib/tools.ts. The order table is fake tutorial data, not a real shop:

TypeScript
import { tool } from 'ai';
import { z } from 'zod';

const orders: Record<string, { status: string; carrier: string; shippedDaysAgo: number; lastScan: string }> = {
  'A-1042': { status: 'in transit', carrier: 'NorthPost', shippedDaysAgo: 9, lastScan: 'Regional sorting hub' },
};

export const lookupOrder = tool({
  description: 'Look up a Northwind Outfitters order by its ID, for example A-1042.',
  inputSchema: z.object({ orderId: z.string().describe('The order ID, for example A-1042') }),
  execute: async ({ orderId }) => orders[orderId.toUpperCase()] ?? { error: 'Order not found' },
});

Replace app/api/chat/route.ts with this complete file. It adds the tool, a line in the instructions that tells the model to use it, and stopWhen: isStepCount(5):

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import {
  convertToModelMessages,
  createUIMessageStreamResponse,
  isStepCount,
  streamText,
  toUIMessageStream,
  type UIMessage,
} from 'ai';
import { lookupOrder } from '../../../lib/tools';

const INSTRUCTIONS = [
  'You are the support assistant for Northwind Outfitters, an outdoor-gear shop.',
  'When a customer mentions an order ID, call the lookupOrder tool before answering.',
  'Keep replies under 120 words and say so plainly when you cannot find an order.',
].join(' ');

export async function POST(req: Request) {
  const { messages }: { messages: UIMessage[] } = await req.json();

  let modelMessages;
  try {
    modelMessages = await convertToModelMessages(messages);
  } catch {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  const result = streamText({
    model: anthropic('claude-sonnet-5-5'),
    instructions: INSTRUCTIONS,
    messages: modelMessages,
    maxOutputTokens: 2048,
    abortSignal: req.signal,
    stopWhen: isStepCount(5),
    tools: { lookupOrder },
  });

  return createUIMessageStreamResponse({
    stream: toUIMessageStream({ stream: result.stream }),
  });
}

A tool call means the model gets a result and then writes its answer, which is a second model call. isStepCount(5) caps how many of those steps one request can take. Five is generous for a single lookup, and the cap stops a confused model from looping on your bill.

On the page, a tool call arrives as a part whose type is tool- plus the tool's name, here tool-lookupOrder. Its state moves from input-streaming and input-available to output-available when the result lands. In app/page.tsx, paste this inside the m.parts.map callback, directly above the return null; line:

TypeScript
            if (part.type === 'tool-lookupOrder') {
              return (
                <p key={i} className="text-sm text-gray-500">
                  {part.state === 'output-available' ? 'Looked up the order' : 'Looking up the order...'}
                </p>
              );
            }

Checkpoint: the dev server reloads on its own. Press Send on the prefilled question again. You should see something like this (an example):

text
You
Where is order A-1042? It has been 9 days.
Assistant
Looked up the order
A-1042 shipped 9 days ago with NorthPost and is still in transit. Its last scan was at a regional sorting hub. I'm sorry about the wait. Want me to open a delivery trace?

If the answer still says it can't look up orders, the route file wasn't saved or the dev server needs a restart.

5. Protect the chat route with a guard

The routes work now, and they're also open to anyone who finds them. A guard module goes in front of them. The rate limiting, input caps and caller identification are tutorial glue, not AI SDK features. Neither Next.js nor the AI SDK provides them.

Create lib/guard.ts. It identifies the caller, rate limits per caller at 10 requests a minute, rejects malformed bodies, and caps a conversation at 30 messages and 20,000 characters:

TypeScript
import type { UIMessage } from 'ai';

const WINDOW_MS = 60_000;
const MAX_REQUESTS_PER_WINDOW = 10;
const MAX_MESSAGES = 30;
const MAX_TOTAL_CHARS = 20_000;
export const MAX_TRIAGE_CHARS = 4_000;

const hits = new Map<string, { count: number; resetAt: number }>();

export function getCallerId(req: Request): string {
  return req.headers.get('x-forwarded-for')?.split(',')[0].trim() || 'local';
}

export function rateLimit(callerId: string): Response | null {
  const now = Date.now();

  if (hits.size > 10_000) {
    for (const [id, entry] of hits) {
      if (entry.resetAt <= now) hits.delete(id);
    }
  }

  const entry = hits.get(callerId);
  if (!entry || entry.resetAt <= now) {
    hits.set(callerId, { count: 1, resetAt: now + WINDOW_MS });
    return null;
  }

  if (entry.count >= MAX_REQUESTS_PER_WINDOW) {
    return Response.json(
      { error: 'Too many requests' },
      { status: 429, headers: { 'retry-after': String(Math.ceil((entry.resetAt - now) / 1000)) } },
    );
  }

  entry.count += 1;
  return null;
}

export async function readJson(req: Request): Promise<unknown> {
  try {
    return await req.json();
  } catch {
    return null;
  }
}

export function validateMessages(body: unknown): UIMessage[] | null {
  if (typeof body !== 'object' || body === null) return null;

  const messages = (body as { messages?: unknown }).messages;
  if (!Array.isArray(messages) || messages.length === 0 || messages.length > MAX_MESSAGES) return null;

  let chars = 0;
  for (const message of messages) {
    if (typeof message !== 'object' || message === null) return null;
    const parts = (message as { parts?: unknown }).parts;
    if (!Array.isArray(parts)) return null;

    for (const part of parts) {
      if (typeof part !== 'object' || part === null) return null;
      if ((part as { type?: unknown }).type === 'text') {
        const text = (part as { text?: unknown }).text;
        if (typeof text !== 'string') return null;
        chars += text.length;
      }
    }
  }

  return chars > MAX_TOTAL_CHARS ? null : (messages as UIMessage[]);
}

Replace app/api/chat/route.ts with this complete file. It calls the guard first and adds three more AI SDK settings: maxRetries, an onError callback that logs the real error on the server, and onEnd, which logs token usage so you can see who's spending what. The toUIMessageStream onError returns the only text the browser sees, so raw provider errors stay on the server:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import {
  APICallError,
  convertToModelMessages,
  createUIMessageStreamResponse,
  isStepCount,
  streamText,
  toUIMessageStream,
} from 'ai';
import { getCallerId, rateLimit, readJson, validateMessages } from '../../../lib/guard';
import { lookupOrder } from '../../../lib/tools';

const INSTRUCTIONS = [
  'You are the support assistant for Northwind Outfitters, an outdoor-gear shop.',
  'When a customer mentions an order ID, call the lookupOrder tool before answering.',
  'Keep replies under 120 words and say so plainly when you cannot find an order.',
].join(' ');

export async function POST(req: Request) {
  const limited = rateLimit(getCallerId(req));
  if (limited) return limited;

  const messages = validateMessages(await readJson(req));
  if (!messages) {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  let modelMessages;
  try {
    modelMessages = await convertToModelMessages(messages);
  } catch {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  const result = streamText({
    model: anthropic('claude-sonnet-5-5'),
    instructions: INSTRUCTIONS,
    messages: modelMessages,
    maxOutputTokens: 2048,
    maxRetries: 2,
    abortSignal: req.signal,
    stopWhen: isStepCount(5),
    tools: { lookupOrder },
    onError: ({ error }) => {
      console.error(error);
    },
    onEnd: (event) => {
      console.log('chat usage', event.usage);
    },
  });

  return createUIMessageStreamResponse({
    stream: toUIMessageStream({
      stream: result.stream,
      onError: (error) =>
        APICallError.isInstance(error) ? 'The AI provider returned an error. Try again.' : 'Something went wrong.',
    }),
  });
}

Two honest limits. The caller ID is the x-forwarded-for header, a stand-in for your real session lookup. When Next.js serves the request itself, it fills the header with the connecting address, so localhost and 127.0.0.1 count as different callers, and any client can send its own value to get a fresh bucket. Behind a proxy that doesn't overwrite the header, it's just as client-controllable. The local fallback only applies when the header is missing entirely. The in-memory limiter also only works within a single server instance, so on a platform that runs several instances, move the counters to Redis or a similar shared store. Real authentication closes both the header gap and the single-instance gap. The page and the route share an origin, so no CORS header is needed. Don't add a wildcard one to silence an error.

Checkpoint: repeat the valid curl -N request from step 2. It should still stream. Then send a malformed body, which must be rejected before any model call happens:

Bash
curl -s -i -X POST http://localhost:3000/api/chat \
  -H "content-type: application/json" \
  -d '{"messages":"nope"}'

You should see something like this (trimmed example):

text
HTTP/1.1 400 Bad Request
content-type: application/json

{"error":"Invalid request"}

If you see Module not found for guard, check that the file is lib/guard.ts, next to app/. If your app has a src/ folder, put both lib/ and app/ inside it.

6. Add a typed triage route

Chat is free-form text. When your code needs to branch on the result, ask for an object instead. A second route calls generateText with output: Output.object({ schema }) and gets back a typed output, as in the AI SDK's structured data guide. Create the schema in lib/schema.ts, the same one from the example at the top:

TypeScript
import { z } from 'zod';

export const triageSchema = z.object({
  category: z.enum(['shipping_delay', 'returns', 'product_question', 'other']),
  urgency: z.enum(['low', 'medium', 'high']),
  summary: z.string(),
});

Then create the route:

Bash
mkdir -p app/api/triage

Create app/api/triage/route.ts. It reuses the guard and takes plain text, which keeps the route easy to test with curl:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import { generateText, NoObjectGeneratedError, Output } from 'ai';
import { MAX_TRIAGE_CHARS, getCallerId, rateLimit, readJson } from '../../../lib/guard';
import { triageSchema } from '../../../lib/schema';

export async function POST(req: Request) {
  const limited = rateLimit(getCallerId(req));
  if (limited) return limited;

  const body = await readJson(req);
  const text = typeof body === 'object' && body !== null ? (body as { text?: unknown }).text : undefined;

  if (typeof text !== 'string' || text.trim() === '') {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }
  if (text.length > MAX_TRIAGE_CHARS) {
    return Response.json({ error: 'Text too long' }, { status: 413 });
  }

  try {
    const { output } = await generateText({
      model: anthropic('claude-sonnet-5-5'),
      output: Output.object({ schema: triageSchema }),
      prompt: `Triage this Northwind Outfitters support conversation.\n\n${text}`,
      maxOutputTokens: 1024,
      abortSignal: req.signal,
    });
    return Response.json(output);
  } catch (error) {
    if (NoObjectGeneratedError.isInstance(error)) {
      console.error('No valid triage object', error.text, error.cause);
      return Response.json({ error: 'The model did not return a valid triage object' }, { status: 502 });
    }
    console.error(error);
    return Response.json({ error: 'Upstream error' }, { status: 502 });
  }
}

The model can still fail to produce an object that matches. When it does, the SDK throws NoObjectGeneratedError, and NoObjectGeneratedError.isInstance(error) identifies it. The error carries the text the model produced and the cause, which is what you log. The route catches it and returns a clean 502 instead of a stack trace.

Checkpoint: in the second terminal, wait a minute if you've just sent several requests, then run:

Bash
curl -X POST http://localhost:3000/api/triage \
  -H "content-type: application/json" \
  -d '{"text":"user: Where is order A-1042? It has been 9 days."}'

You should see something like this (an example, the summary wording varies):

text
{"category":"shipping_delay","urgency":"high","summary":"Customer asks about order A-1042, shipped 9 days ago and still in transit."}

An empty body is rejected the same way as in step 5:

Bash
curl -i -X POST http://localhost:3000/api/triage \
  -H "content-type: application/json" \
  -d '{"text":""}'

That returns a 400 with {"error":"Invalid request"}.

7. Add the triage button to the page

The page needs a button that sends the conversation to /api/triage and shows the JSON. Replace app/page.tsx with this complete file. It's the step 3 page plus the tool part from step 4 plus a runTriage function, a button and a result box. The fetch code is tutorial glue:

TypeScript
'use client';

import { useChat } from '@ai-sdk/react';
import { DefaultChatTransport } from 'ai';
import { useState } from 'react';

const transport = new DefaultChatTransport({ api: '/api/chat' });

export default function Page() {
  const [input, setInput] = useState('Where is order A-1042? It has been 9 days.');
  const [triage, setTriage] = useState('');
  const { messages, sendMessage, status, error, regenerate, stop } = useChat({ transport });

  const busy = status === 'submitted' || status === 'streaming';

  async function runTriage() {
    const text = messages
      .map((m) => `${m.role}: ${m.parts.map((p) => (p.type === 'text' ? p.text : '')).join('')}`)
      .join('\n');

    try {
      const res = await fetch('/api/triage', {
        method: 'POST',
        headers: { 'content-type': 'application/json' },
        body: JSON.stringify({ text }),
      });
      setTriage(JSON.stringify(await res.json(), null, 2));
    } catch {
      setTriage('Triage request failed');
    }
  }

  return (
    <main className="mx-auto max-w-2xl space-y-4 p-6">
      <h1 className="text-xl font-semibold">Northwind Outfitters support</h1>

      {messages.map((m) => (
        <div key={m.id}>
          <strong>{m.role === 'user' ? 'You' : 'Assistant'}</strong>
          {m.parts.map((part, i) => {
            if (part.type === 'text') {
              return (
                <p key={i} className="whitespace-pre-wrap">
                  {part.text}
                </p>
              );
            }
            if (part.type === 'tool-lookupOrder') {
              return (
                <p key={i} className="text-sm text-gray-500">
                  {part.state === 'output-available' ? 'Looked up the order' : 'Looking up the order...'}
                </p>
              );
            }
            return null;
          })}
        </div>
      ))}

      {error && (
        <p className="text-red-600">
          {error.message}{' '}
          <button type="button" className="underline" onClick={() => regenerate()}>
            Retry
          </button>
        </p>
      )}

      <form
        className="flex gap-2"
        onSubmit={(e) => {
          e.preventDefault();
          if (busy || input.trim() === '') return;
          sendMessage({ text: input });
          setInput('');
        }}
      >
        <input
          className="flex-1 border p-2"
          value={input}
          onChange={(e) => setInput(e.currentTarget.value)}
          placeholder="Ask about an order"
        />
        <button type="submit" className="border px-3" disabled={busy}>
          Send
        </button>
        <button type="button" className="border px-3" onClick={() => stop()} disabled={!busy}>
          Stop
        </button>
      </form>

      <button
        type="button"
        className="border px-3 py-1"
        onClick={runTriage}
        disabled={busy || messages.length === 0}
      >
        Triage this conversation
      </button>
      {triage && <pre className="bg-gray-100 p-3 text-sm">{triage}</pre>}
    </main>
  );
}

Checkpoint: in the browser, send the prefilled question, wait for the answer, then click "Triage this conversation". You should see something like this under the button (an example):

text
{
  "category": "shipping_delay",
  "urgency": "high",
  "summary": "Customer asks about order A-1042, shipped 9 days ago and still in transit."
}

8. Run it end to end and test the failures

Your files now match the Full code section. Check them with the type checker, from northwind-support/:

Bash
npx tsc --noEmit

It prints nothing when the types check. Then, in the browser at http://localhost:3000, run the whole feature. Send the prefilled question, click the triage button, and compare the results with the transcript and JSON in What you'll build. The terminal running npm run dev should also show a line starting with chat usage, followed by the token usage for that request. If no chat usage line appears and you see AI_APICallError instead, the request failed before the model answered, so the key is wrong. Check .env.local.

Now a negative check. Replace the input text with Where is order Z-9999? and send it. The table has no such order, so the tool returns an error object and the assistant should say plainly that it can't find the order (example, wording varies):

text
You
Where is order Z-9999?
Assistant
Looked up the order
I couldn't find order Z-9999. Could you double-check the ID?

Last, test the rate limit. Wait 60 seconds so the window resets, then send 12 malformed requests in a row from the second terminal. They never reach the model:

Bash
for i in $(seq 1 12); do curl -s -o /dev/null -w "%{http_code}\n" -X POST http://localhost:3000/api/chat -H "content-type: application/json" -d '{"messages":"nope"}'; done

You should see ten 400 responses, then two 429 responses:

text
400
400
400
400
400
400
400
400
400
400
429
429

If you see a 429 earlier, you sent requests in the same minute before the loop. Wait 60 seconds and run it again.

Architecture

text
browser (useChat, no key in the page)
  |  POST /api/chat  { messages }
  v
/api/chat route: rate limit, body checks, message and character caps
  |  streamText + lookupOrder tool (ANTHROPIC_API_KEY lives only here)
  v
Anthropic API (claude-sonnet-5-5)
  |  model output
  v
/api/chat route: toUIMessageStream -> createUIMessageStreamResponse
  |  server-sent events (UI message stream)
  v
browser (useChat renders message.parts)

browser (Triage button)
  |  POST /api/triage  { text }
  v
/api/triage route: same guards -> generateText + Output.object -> JSON

Clean up

Stop the dev server with Ctrl+C in its terminal. To remove the project, run rm -rf northwind-support from the folder that contains it. Then delete the API key you created for this tutorial in the Claude Console, unless you use it elsewhere.

Full code

These are the files your steps produce. Use this section to compare or to copy a file you got wrong. Package versions are pinned to the majors checked on 2026-10-01: ai 7, @ai-sdk/react 4, @ai-sdk/anthropic 4 and zod 4. The routes and the page use only the APIs described in the steps, and the glue (guard, page layout) is tutorial code.

The project layout, with lib/ next to app/ (if your app has a src/ folder, put both inside it):

text
northwind-support/
  app/
    api/
      chat/route.ts
      triage/route.ts
    page.tsx
  lib/
    guard.ts
    schema.ts
    tools.ts
  .env.local
  .env.example
  package.json

Install and run:

Bash
npx create-next-app@latest northwind-support --yes
cd northwind-support
printf '\n!.env.example\n' >> .gitignore
npm install ai@7 @ai-sdk/react@4 @ai-sdk/anthropic@4 zod@4
cp .env.example .env.local   # then put your real key in .env.local
npm run dev

.env.example:

Bash
ANTHROPIC_API_KEY=

package.json gets four new dependencies. next, react and react-dom come from create-next-app and stay as generated. The new entries are the pins that matter:

JSON
{
  "dependencies": {
    "@ai-sdk/anthropic": "^4.0.71",
    "@ai-sdk/react": "^4.0.129",
    "ai": "^7.0.126",
    "zod": "^4.6.5"
  }
}

lib/schema.ts:

TypeScript
import { z } from 'zod';

export const triageSchema = z.object({
  category: z.enum(['shipping_delay', 'returns', 'product_question', 'other']),
  urgency: z.enum(['low', 'medium', 'high']),
  summary: z.string(),
});

lib/tools.ts:

TypeScript
import { tool } from 'ai';
import { z } from 'zod';

const orders: Record<string, { status: string; carrier: string; shippedDaysAgo: number; lastScan: string }> = {
  'A-1042': { status: 'in transit', carrier: 'NorthPost', shippedDaysAgo: 9, lastScan: 'Regional sorting hub' },
};

export const lookupOrder = tool({
  description: 'Look up a Northwind Outfitters order by its ID, for example A-1042.',
  inputSchema: z.object({ orderId: z.string().describe('The order ID, for example A-1042') }),
  execute: async ({ orderId }) => orders[orderId.toUpperCase()] ?? { error: 'Order not found' },
});

lib/guard.ts (tutorial glue: swap getCallerId for your real session lookup, and replace the in-memory Map with a shared store if you run more than one instance). When Next.js serves the request itself, it fills x-forwarded-for with the connecting address, so localhost and 127.0.0.1 count as different callers, and any client can send its own value to get a fresh bucket. Behind a proxy that doesn't overwrite the header, it's just as client-controllable. The local fallback only applies when the header is missing entirely. Real authentication closes the gap:

TypeScript
import type { UIMessage } from 'ai';

const WINDOW_MS = 60_000;
const MAX_REQUESTS_PER_WINDOW = 10;
const MAX_MESSAGES = 30;
const MAX_TOTAL_CHARS = 20_000;
export const MAX_TRIAGE_CHARS = 4_000;

const hits = new Map<string, { count: number; resetAt: number }>();

export function getCallerId(req: Request): string {
  return req.headers.get('x-forwarded-for')?.split(',')[0].trim() || 'local';
}

export function rateLimit(callerId: string): Response | null {
  const now = Date.now();

  if (hits.size > 10_000) {
    for (const [id, entry] of hits) {
      if (entry.resetAt <= now) hits.delete(id);
    }
  }

  const entry = hits.get(callerId);
  if (!entry || entry.resetAt <= now) {
    hits.set(callerId, { count: 1, resetAt: now + WINDOW_MS });
    return null;
  }

  if (entry.count >= MAX_REQUESTS_PER_WINDOW) {
    return Response.json(
      { error: 'Too many requests' },
      { status: 429, headers: { 'retry-after': String(Math.ceil((entry.resetAt - now) / 1000)) } },
    );
  }

  entry.count += 1;
  return null;
}

export async function readJson(req: Request): Promise<unknown> {
  try {
    return await req.json();
  } catch {
    return null;
  }
}

export function validateMessages(body: unknown): UIMessage[] | null {
  if (typeof body !== 'object' || body === null) return null;

  const messages = (body as { messages?: unknown }).messages;
  if (!Array.isArray(messages) || messages.length === 0 || messages.length > MAX_MESSAGES) return null;

  let chars = 0;
  for (const message of messages) {
    if (typeof message !== 'object' || message === null) return null;
    const parts = (message as { parts?: unknown }).parts;
    if (!Array.isArray(parts)) return null;

    for (const part of parts) {
      if (typeof part !== 'object' || part === null) return null;
      if ((part as { type?: unknown }).type === 'text') {
        const text = (part as { text?: unknown }).text;
        if (typeof text !== 'string') return null;
        chars += text.length;
      }
    }
  }

  return chars > MAX_TOTAL_CHARS ? null : (messages as UIMessage[]);
}

app/api/chat/route.ts:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import {
  APICallError,
  convertToModelMessages,
  createUIMessageStreamResponse,
  isStepCount,
  streamText,
  toUIMessageStream,
} from 'ai';
import { getCallerId, rateLimit, readJson, validateMessages } from '../../../lib/guard';
import { lookupOrder } from '../../../lib/tools';

const INSTRUCTIONS = [
  'You are the support assistant for Northwind Outfitters, an outdoor-gear shop.',
  'When a customer mentions an order ID, call the lookupOrder tool before answering.',
  'Keep replies under 120 words and say so plainly when you cannot find an order.',
].join(' ');

export async function POST(req: Request) {
  const limited = rateLimit(getCallerId(req));
  if (limited) return limited;

  const messages = validateMessages(await readJson(req));
  if (!messages) {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  let modelMessages;
  try {
    modelMessages = await convertToModelMessages(messages);
  } catch {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }

  const result = streamText({
    model: anthropic('claude-sonnet-5-5'),
    instructions: INSTRUCTIONS,
    messages: modelMessages,
    maxOutputTokens: 2048,
    maxRetries: 2,
    abortSignal: req.signal,
    stopWhen: isStepCount(5),
    tools: { lookupOrder },
    onError: ({ error }) => {
      console.error(error);
    },
    onEnd: (event) => {
      console.log('chat usage', event.usage);
    },
  });

  return createUIMessageStreamResponse({
    stream: toUIMessageStream({
      stream: result.stream,
      onError: (error) =>
        APICallError.isInstance(error) ? 'The AI provider returned an error. Try again.' : 'Something went wrong.',
    }),
  });
}

app/api/triage/route.ts:

TypeScript
import { anthropic } from '@ai-sdk/anthropic';
import { generateText, NoObjectGeneratedError, Output } from 'ai';
import { MAX_TRIAGE_CHARS, getCallerId, rateLimit, readJson } from '../../../lib/guard';
import { triageSchema } from '../../../lib/schema';

export async function POST(req: Request) {
  const limited = rateLimit(getCallerId(req));
  if (limited) return limited;

  const body = await readJson(req);
  const text = typeof body === 'object' && body !== null ? (body as { text?: unknown }).text : undefined;

  if (typeof text !== 'string' || text.trim() === '') {
    return Response.json({ error: 'Invalid request' }, { status: 400 });
  }
  if (text.length > MAX_TRIAGE_CHARS) {
    return Response.json({ error: 'Text too long' }, { status: 413 });
  }

  try {
    const { output } = await generateText({
      model: anthropic('claude-sonnet-5-5'),
      output: Output.object({ schema: triageSchema }),
      prompt: `Triage this Northwind Outfitters support conversation.\n\n${text}`,
      maxOutputTokens: 1024,
      abortSignal: req.signal,
    });
    return Response.json(output);
  } catch (error) {
    if (NoObjectGeneratedError.isInstance(error)) {
      console.error('No valid triage object', error.text, error.cause);
      return Response.json({ error: 'The model did not return a valid triage object' }, { status: 502 });
    }
    console.error(error);
    return Response.json({ error: 'Upstream error' }, { status: 502 });
  }
}

app/page.tsx (replaces the generated page; the layout stays as generated):

TypeScript
'use client';

import { useChat } from '@ai-sdk/react';
import { DefaultChatTransport } from 'ai';
import { useState } from 'react';

const transport = new DefaultChatTransport({ api: '/api/chat' });

export default function Page() {
  const [input, setInput] = useState('Where is order A-1042? It has been 9 days.');
  const [triage, setTriage] = useState('');
  const { messages, sendMessage, status, error, regenerate, stop } = useChat({ transport });

  const busy = status === 'submitted' || status === 'streaming';

  async function runTriage() {
    const text = messages
      .map((m) => `${m.role}: ${m.parts.map((p) => (p.type === 'text' ? p.text : '')).join('')}`)
      .join('\n');

    try {
      const res = await fetch('/api/triage', {
        method: 'POST',
        headers: { 'content-type': 'application/json' },
        body: JSON.stringify({ text }),
      });
      setTriage(JSON.stringify(await res.json(), null, 2));
    } catch {
      setTriage('Triage request failed');
    }
  }

  return (
    <main className="mx-auto max-w-2xl space-y-4 p-6">
      <h1 className="text-xl font-semibold">Northwind Outfitters support</h1>

      {messages.map((m) => (
        <div key={m.id}>
          <strong>{m.role === 'user' ? 'You' : 'Assistant'}</strong>
          {m.parts.map((part, i) => {
            if (part.type === 'text') {
              return (
                <p key={i} className="whitespace-pre-wrap">
                  {part.text}
                </p>
              );
            }
            if (part.type === 'tool-lookupOrder') {
              return (
                <p key={i} className="text-sm text-gray-500">
                  {part.state === 'output-available' ? 'Looked up the order' : 'Looking up the order...'}
                </p>
              );
            }
            return null;
          })}
        </div>
      ))}

      {error && (
        <p className="text-red-600">
          {error.message}{' '}
          <button type="button" className="underline" onClick={() => regenerate()}>
            Retry
          </button>
        </p>
      )}

      <form
        className="flex gap-2"
        onSubmit={(e) => {
          e.preventDefault();
          if (busy || input.trim() === '') return;
          sendMessage({ text: input });
          setInput('');
        }}
      >
        <input
          className="flex-1 border p-2"
          value={input}
          onChange={(e) => setInput(e.currentTarget.value)}
          placeholder="Ask about an order"
        />
        <button type="submit" className="border px-3" disabled={busy}>
          Send
        </button>
        <button type="button" className="border px-3" onClick={() => stop()} disabled={!busy}>
          Stop
        </button>
      </form>

      <button
        type="button"
        className="border px-3 py-1"
        onClick={runTriage}
        disabled={busy || messages.length === 0}
      >
        Triage this conversation
      </button>
      {triage && <pre className="bg-gray-100 p-3 text-sm">{triage}</pre>}
    </main>
  );
}

Open http://localhost:3000, send the prefilled question, then click the triage button. Answers need a real key in .env.local.

Reference

Migrating from v6

If you're porting older code, the AI SDK 7.0 migration guide lists these renames. In AI SDK 7.0.126 result.toUIMessageStreamResponse() still exists, but its type definition marks it @deprecated and says it will go in the next major release.

v6v7
result.toUIMessageStreamResponse()createUIMessageStreamResponse({ stream: toUIMessageStream({ stream: result.stream }) })
stepCountIs(n)isStepCount(n)
systeminstructions
experimental_outputoutput
onFinishonEnd

The guide also says v7 is ESM only and ships a codemod: npx @ai-sdk/codemod v7.

Retries and errors

The AI SDK's error handling page describes the channels failures arrive through. In streamText, errors reach an onError callback and show up as error parts in the stream. maxRetries and streamRetries have different defaults.

Where it failsMechanismWhat to do
Starting the callmaxRetries, default 2Leave it. Retries cover transient failures.
After streaming beginsstreamRetries, default 0Leave it off for chat unless you've tested how a restarted stream renders.
Inside the streamonError and error partsLog the real error server-side, forward a safe string to the browser.
Bad key or malformed requestSame failure every timeDon't retry. Fix the key or the request.

StreamProviderError carries type, code, statusCode and isRetryable, so a custom retry policy can check isRetryable instead of guessing from status codes. Failures that begin mid-stream can arrive as a StreamProviderError, which the step 5 mapper reports as the generic message. Thinking tokens count toward maxOutputTokens when thinking is enabled, so a cap that's too low can cut a reply off before the visible answer ends.

Cost

Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, per the models overview. At those prices, one chat turn with 2,000 input tokens and 400 output tokens is $0.004 plus $0.004, or $0.008, before thinking tokens. That's arithmetic from the list prices, not a measurement.

Three things move the number. The output cap sets the ceiling per reply. The conversation length matters because the whole history goes in with every request, and a tool step adds another model call that resends it, which is why the guard caps messages and characters. And long fixed prefixes, such as a big instructions block, can be cached: the provider docs list a cacheControl option under providerOptions.anthropic. They show the breakpoint as providerOptions: { anthropic: { cacheControl: { type: 'ephemeral' } } } on the message or part you want cached. This tutorial's prompts are short, so the code doesn't include it.

Switching provider

Changing the model is a small edit, and it's one of the main reasons teams pick the AI SDK over a single vendor SDK. For OpenAI, install the provider and set OPENAI_API_KEY:

Bash
npm install @ai-sdk/openai@4

For illustration only, not part of this project, swap the model expression in streamText:

TypeScript
import { openai } from '@ai-sdk/openai';

// in streamText, replace the model line
model: openai('gpt-5.6'),

The OpenAI provider page lists the model IDs. There's also the Vercel AI Gateway route: set AI_GATEWAY_API_KEY and pass a plain string such as "anthropic/claude-sonnet-5.5" as the model. You'd use that if you want one key and one bill across providers. This tutorial uses the direct provider, so your Anthropic key and usage stay under your own account.

Deploying

Set ANTHROPIC_API_KEY in your host's environment variable settings. .env.local stays on your machine. You don't need Vercel hosting for any of this. Any host that runs a Node.js server for Next.js works.

Common mistakes

  • Copying v4 to v6 code. toDataStreamResponse(), handleSubmit, input from useChat and maxSteps are gone or deprecated in AI SDK 7. The rename table in Reference maps each one, and steps 2 and 3 use the v7 forms.
  • Naming the key NEXT_PUBLIC_ANTHROPIC_API_KEY. Next.js bundles every NEXT_PUBLIC_ variable into the browser, so the key becomes public (step 1).
  • Assuming the first message part is text. Reasoning parts and tool parts can arrive first, so branch on part.type (step 3).
  • Shipping without maxOutputTokens. One long reply, or one loop of tool steps, sets your cost instead of you (steps 2 and 4).
  • Skipping abortSignal: req.signal. The model keeps generating after the user closes the tab, and you pay for it (step 2).
  • Opening the route to every origin with a wildcard CORS header to silence a cross-origin error. The page and the route share an origin here. Keep it that way (step 5).

Next steps

  • Persist conversations and replace getCallerId with real authentication. The route currently trusts the browser to send the history, so store messages server-side, load them by conversation ID, and move the rate limiter to a shared store.
  • Add more tools. The same tool() shape works for refunds, stock checks or an MCP server, with isStepCount still capping the loop.
  • Try the alternatives kept out of the steps: the provider swap and gateway route in Reference, and streamText with partialOutputStream to show the triage object as it fills in. The second needs extra framework code to get the partials to the browser, which this tutorial skips because the triage object is small and arrives fast.
  • Build a small eval set of 20 real support questions and run it whenever you change the instructions or the model ID. For the Messages API underneath all of this, see calling Claude directly with Anthropic's own SDK.

Hiring React and Node.js engineers through HighCircl

If you'd rather add an engineer than build this alone, HighCircl covers React and Node.js, among other stacks, in seven European countries. Senior engineers run four vetting stages, and about 1 in 10 applicants passes. You get a shortlist of 3-5 candidates within 72 hours. Rates run €45-105/hr ($50-115/hr), with a transparent 20% capped margin, no subscription and no minimum hours, plus a replacement guarantee. See HighCircl React developers.

FAQ

Do I need Vercel hosting to use the Vercel AI SDK?

No. The AI SDK is an open library, and the routes in this tutorial run on any host that serves Next.js on Node.js 22 or later. Vercel's own platform and AI Gateway are optional. Set ANTHROPIC_API_KEY in your host's environment settings and deploy.

Should I use the AI Gateway or the direct Anthropic provider?

This tutorial uses the direct provider, @ai-sdk/anthropic with ANTHROPIC_API_KEY. The getting-started guide uses the gateway instead: you set AI_GATEWAY_API_KEY and pass a string like "anthropic/claude-sonnet-5.5". The gateway gives you one key across providers. The direct provider keeps your usage and billing in your own Anthropic account.

Can I use OpenAI instead of Claude?

Yes. Install @ai-sdk/openai, set OPENAI_API_KEY, and change the model line to openai('gpt-5.6'). Nothing else in the routes or the page changes, which is the point of the provider layer. Tool and structured output behaviour can differ between models, so re-run your tests after the swap.

Is my API key safe with useChat?

Yes, as long as the key stays in a server environment variable without the NEXT_PUBLIC_ prefix. useChat runs in the browser but only calls your own /api/chat route. The route holds the key and talks to Anthropic. Add a rate limit and input caps to that route, as in step 5, so nobody can use your endpoint to spend your money.

How do I cap the cost of an AI SDK chat?

Set maxOutputTokens on every call, cap the number of tool steps with isStepCount, limit how long a conversation can get, and abort on disconnect with abortSignal. Sonnet 5.5 is $2 per million input tokens and $10 per million output tokens. Log usage from onEnd so you can see where the spend goes.

What changed in AI SDK 7?

Several names. stepCountIs became isStepCount, system became instructions, onFinish became onEnd, and experimental_output became output. streamText results now go through createUIMessageStreamResponse with toUIMessageStream, the package is ESM only, and Node.js 22 or later is required. A codemod, npx @ai-sdk/codemod v7, handles most of the mechanical work.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now