October 1, 2026

How to use the Claude API: build an AI feature step by step

Wire the Claude API into your own project: get a key, make your first request, add a backend endpoint, streaming, retries and cost limits, with code.

Guide

Tech

This is how to use the Claude API in a real project, not in a throwaway script. You'll end up with a feedback triage feature that returns JSON your code can trust, streams a summary into the browser as it's written, and never lets the API key leave your server. If you're handing keys to a team, AI in the engineering workflow: tools, permissions and policy covers who should get them and under what rules.

What you'll build

A feedback triage feature: your backend takes a block of customer text, returns {category, urgency, summary}, and streams a two-sentence summary to a small web page. Plan on about an hour, less if you already have a key. The steps use Python with FastAPI. A complete TypeScript and Express version is in Full code.

Here's the finished endpoint, called with curl. The values in the response are an example of the shape, not real output:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'
JSON
{
  "category": "bug",
  "urgency": "high",
  "summary": "Customer reports the invoice page crashes when they click download."
}

What you need before you start

  • Python 3.10 or later, the minimum the Python SDK page states. For the TypeScript version in Full code, use Node.js 22 LTS or later (Node 20 reached end of life in April 2026) and TypeScript 5.0 or later.
  • A macOS or Linux shell. On Windows, use WSL.
  • A Claude Console account and an API key. The click path, as Anthropic's API key guide describes it:
    1. Go to platform.claude.com and sign in, or create an account.
    2. Open Settings > API keys. The direct address is https://platform.claude.com/settings/keys.
    3. Click Create key.
    4. Give the key a name and choose an expiration.
    5. Set Linked account to yourself or to a service account. You can optionally pick a workspace too.

Copy the key straight away. The guide says the Console shows the full key, which starts with sk-ant-, only once, at creation, and that if you lose it you can't view it again: you create a new one instead. If Create key is greyed out, your role may not allow you to create keys there, and an organization admin needs to do it for you.

Which linked account? The guide's rule is short: use a personal key for your own development, and a service account key for anything shared. Your laptop experiments get a personal key. A deployed backend should run on a service account key.

On cost, the pricing page's FAQ says new users receive a small amount of free credits to test the API. Billing lives under Settings > Billing, which you'll visit in step 7.

Quickstart

Already have a key and just want a first call? Install the SDK, export the key, and run this:

Bash
pip install anthropic
export ANTHROPIC_API_KEY="your-api-key-here"
Python
import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1000,
    messages=[{"role": "user", "content": "Summarise this feedback in one sentence: The invoice page crashes when I click download."}],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

Three things to get right before you build on it:

  • Keep the key on the server. Never put it in frontend code or source control.
  • Set a spend limit in the Console under Settings > Billing, so a bug can't run up an open-ended bill.
  • Loop over the content blocks and check block.type. Don't read content[0].text, because the first block can be a thinking block.

How to use the Claude API in your own project

1. Create the project and store your key

Make a project folder with a virtual environment, install every package the tutorial uses, and put your key in a git-ignored .env file. Everything later builds on this folder.

In a terminal, from wherever you keep projects:

Bash
mkdir claude-quickstart && cd claude-quickstart
python3 -m venv .venv && source .venv/bin/activate
pip install anthropic fastapi uvicorn python-dotenv
mkdir static

Create .env.example, a template that's safe to commit:

text
ANTHROPIC_API_KEY=

Copy it to .env and paste your real key after the equals sign:

Bash
cp .env.example .env

Create .gitignore so the key and the environment never reach git:

text
.env
.venv/
__pycache__/

ANTHROPIC_API_KEY is the variable name the quickstart uses, and the SDK reads it automatically from the shell. The SDKs don't read .env files, so every Python file in this tutorial calls load_dotenv() from python-dotenv before it creates the client. A key that lands in git history is a leaked key, even if you delete the file in the next commit.

Checkpoint, in the same terminal:

Bash
pip show anthropic

You should see something like (the version number will differ):

text
Name: anthropic
Version: <installed version>

If your prompt doesn't start with (.venv), re-run the source .venv/bin/activate line.

2. Send your first request

Create quickstart.py next to .env. It sends one message to the model and prints the text blocks in the reply, which proves your key, install and network path work.

Python
import anthropic
from dotenv import load_dotenv

load_dotenv()

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1000,
    messages=[
        {
            "role": "user",
            "content": "Summarise this customer feedback in one sentence: The invoice page crashes when I click download.",
        }
    ],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

Run it:

Bash
python quickstart.py

You should see something like (the wording will differ):

text
A customer reports that the invoice page crashes whenever they click the download button.

If you see AuthenticationError, the key isn't loaded: check that .env sits in this folder and holds your real key.

claude-sonnet-5-5 is the current Sonnet from Anthropic's models overview. Many tutorials still teach claude-3-* IDs, which are stale.

Why loop over message.content instead of reading message.content[0].text? The Sonnet 5.5 overview says adaptive thinking is on by default, and the thinking page says that up-front thinking arrives in thinking content blocks ahead of the response. The first block can have no .text at all, so code that assumes index zero will crash on some requests and work on others. Also don't set temperature, top_p or top_k to anything but the default: on Sonnet 5.5 that returns a 400 error.

3. Put the call behind your own backend endpoint

Create main.py with a FastAPI endpoint that calls Claude on the server. Your browser or app calls this endpoint, and only this endpoint holds the key. That gives you one place for auth, input limits and rate limiting.

text
browser
  |  POST /api/triage (no key in the page)
  v
your endpoint: auth, input length cap, per-user rate limit
  |  x-api-key (lives only on the server)
  v
Claude API

The FastAPI routes and the browser page later on are web-framework glue, not Anthropic SDK calls. They're minimal code written for this tutorial. Create main.py:

Python
import anthropic
from dotenv import load_dotenv
from fastapi import FastAPI
from pydantic import BaseModel

load_dotenv()

app = FastAPI()
client = anthropic.Anthropic()


class Feedback(BaseModel):
    text: str


@app.post("/api/triage")
def triage(body: Feedback):
    message = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1000,
        messages=[{"role": "user", "content": body.text[:4000]}],
    )
    return {"text": next((b.text for b in message.content if b.type == "text"), "")}

Start the server in one terminal, from claude-quickstart/ with the virtual environment active:

Bash
uvicorn main:app --reload

In a second terminal, send a request:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (the wording will differ):

JSON
{"text":"Thanks for flagging this. A crash on the invoice download needs a bug report with the browser and steps to reproduce."}

If uvicorn says the address is already in use, stop the other process or add --port 8001 and change the curl URL to match.

The body.text[:4000] slice is the input length cap. Without one, anyone who finds your endpoint can send a 10 MB document and you pay for it. This endpoint has no authentication yet, which is fine on localhost and not fine anywhere else, so put your session or token check in front of it before you deploy.

The reason for the proxy: the TypeScript SDK page says browser use is disabled by default to avoid exposing your secret API credentials. An option named dangerouslyAllowBrowser overrides that, and the same page warns that any user with access to the browser can potentially inspect, extract and misuse those credentials. A key in frontend code is a key you've published. What to check before a vibe coded app reaches real users covers how that failure shows up in shipped apps.

4. Return structured output

Replace main.py so the endpoint returns a typed {category, urgency, summary} object instead of free text. You define the shape once as a Pydantic model, call messages.parse, and read the validated object from parsed_output.

Instructions that apply from the first turn belong in the top-level system field, per the working with messages guide. Here the triage instruction travels in the user message instead, because the documented parse examples don't combine it with system. Prefilling the assistant's reply, the old trick for forcing JSON, is gone: on Claude 4.6 and later a request with prefill returns a 400, and the docs point you to structured outputs. claude-sonnet-5-5 is on the structured outputs page's supported list, and the feature is generally available.

Replace the whole of main.py with:

Python
from typing import Literal

import anthropic
from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, ValidationError

load_dotenv()

app = FastAPI()
client = anthropic.Anthropic()


class Feedback(BaseModel):
    text: str


class Triage(BaseModel):
    category: Literal["bug", "billing", "feature_request", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str


@app.post("/api/triage")
def triage(body: Feedback):
    try:
        response = client.messages.parse(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[{"role": "user", "content": "Triage this customer feedback. Be concise.\n\n" + body.text[:4000]}],
            output_format=Triage,
        )
    except ValidationError:
        raise HTTPException(status_code=502, detail="No usable result (output did not match the schema)")
    if response.stop_reason in ("max_tokens", "refusal") or response.parsed_output is None:
        raise HTTPException(status_code=502, detail=f"No usable result ({response.stop_reason}), request {response._request_id}")
    return response.parsed_output

The if is the guard that matters. A successful HTTP response isn't a successful answer: end_turn is normal, max_tokens means the reply was cut off, and refusal means the model declined. Refusals have billing consequences that depend on the model, covered in how refusals are billed before any output on some models. When the reply is cut off or the refusal is written as text, parse can't validate it and raises ValidationError before you see stop_reason, so catch that too. With both in place, every unusable result becomes a 502 you control.

With --reload running, the server restarts on save. Run the same curl as in step 3.

You should see something like (the values will differ):

JSON
{"category":"bug","urgency":"high","summary":"Customer reports the invoice page crashes when they click download."}

5. Stream a summary to the browser

Add a streaming endpoint to main.py and a page at static/index.html that shows text as it arrives. Streaming sends the reply to the user as the model writes it, so a long answer starts appearing early.

The streaming page documents the SDK's stream helper, which the new route uses. Under the hood, "stream": true on the raw API makes it respond with server-sent events. The Messages API docs also recommend streaming for long-running requests. Replace the whole of main.py with:

Python
import json
from typing import Literal

import anthropic
from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from fastapi.staticfiles import StaticFiles
from pydantic import BaseModel, ValidationError

load_dotenv()

app = FastAPI()
client = anthropic.Anthropic()


class Feedback(BaseModel):
    text: str


class Triage(BaseModel):
    category: Literal["bug", "billing", "feature_request", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str


@app.post("/api/triage")
def triage(body: Feedback):
    try:
        response = client.messages.parse(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[{"role": "user", "content": "Triage this customer feedback. Be concise.\n\n" + body.text[:4000]}],
            output_format=Triage,
        )
    except ValidationError:
        raise HTTPException(status_code=502, detail="No usable result (output did not match the schema)")
    if response.stop_reason in ("max_tokens", "refusal") or response.parsed_output is None:
        raise HTTPException(status_code=502, detail=f"No usable result ({response.stop_reason}), request {response._request_id}")
    return response.parsed_output


@app.post("/api/triage/stream")
def triage_stream(body: Feedback):
    def generate():
        with client.messages.stream(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[
                {
                    "role": "user",
                    "content": "Summarise this feedback in two sentences:\n\n" + body.text[:4000],
                }
            ],
        ) as stream:
            for text in stream.text_stream:
                yield f"data: {json.dumps(text)}\n\n"

    return StreamingResponse(generate(), media_type="text/event-stream")


app.mount("/", StaticFiles(directory="static", html=True), name="static")

The static mount sits last on purpose. A mount at / matches everything, so it would swallow any route registered after it. Each chunk is JSON-encoded so newlines inside the text can't break the event framing.

Now the page. A plain fetch with a stream reader does the job, because EventSource only sends GET requests and you're posting a body. Create static/index.html:

html
<!doctype html>
<html>
  <body>
    <textarea id="feedback" rows="6" cols="60"></textarea>
    <button id="go">Summarise</button>
    <div id="output" style="white-space: pre-wrap"></div>

    <script>
      const output = document.getElementById("output");

      document.getElementById("go").addEventListener("click", async () => {
        const text = document.getElementById("feedback").value;
        output.textContent = "";

        const res = await fetch("/api/triage/stream", {
          method: "POST",
          headers: { "content-type": "application/json" },
          body: JSON.stringify({ text })
        });

        const reader = res.body.getReader();
        const decoder = new TextDecoder();
        let buffer = "";

        while (true) {
          const { done, value } = await reader.read();
          if (done) break;
          buffer += decoder.decode(value, { stream: true });
          const events = buffer.split("\n\n");
          buffer = events.pop() ?? "";
          for (const event of events) {
            output.textContent += JSON.parse(event.slice("data: ".length));
          }
        }
      });
    </script>
  </body>
</html>

Checkpoint one, in the second terminal. The -N flag turns off curl's buffering so you see chunks as they arrive:

Bash
curl -N -X POST http://localhost:8000/api/triage/stream \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (the chunk boundaries will differ):

text
data: "A customer reports"

data: " that the invoice page crashes"

data: " on download."

Checkpoint two: open http://localhost:8000, paste the same sentence into the box and click Summarise. The summary should appear word by word in the div below the button.

The relative fetch URL works only when the page comes from the same origin as the API, as it does here. Don't "fix" a cross-origin error by opening the API to every origin with allow_origins=["*"]. And per the errors documentation, an error can occur after the API returns a 200 response, so a stream that began fine can still stop partway. Your UI needs a state for "this stopped early". The same applies if the API call itself fails, for example with a wrong key: the route has already sent its 200 headers, so the page shows an empty output div and the error appears only in the uvicorn log. If you see an empty response, check that log first.

6. Trigger an API error and read it

Break the key on purpose to see what an SDK error looks like from your endpoint, before a real failure teaches you. This step changes no files.

Stop uvicorn with Ctrl+C and restart it with a wrong key. A variable set in the shell takes precedence over .env, because load_dotenv() doesn't override existing variables by default:

Bash
ANTHROPIC_API_KEY=wrong-key uvicorn main:app

In the second terminal, run the curl from step 3 again. You should see something like:

text
Internal Server Error

In the uvicorn terminal, the traceback should end with a line like this (the message and request ID will differ):

text
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': '<message text>'}, 'request_id': 'req_<id>'}

The errors page lists the HTTP codes and error types, and every error body carries a request_id. The same ID comes back in the request-id response header, and the SDKs expose it as _request_id. Log it on every failure, so you can trace a bad response later.

You didn't write retry code, and a 401 wouldn't have used it anyway. The SDK retries transient failures for you, with exponential backoff, twice by default, honoring the retry-after header when present. It retries connection errors, 408, 409, 429 and 500 and above.

There are two different 429s, and the tutorials we checked skip that. A normal rate-limit 429 carries a retry-after header with the seconds to wait, so waiting works. A spend-cap 429 has error.details.error_code set to enforced_spend_limit_reached and no retry-after header, and the rate limits page is blunt that retrying, including the SDK's automatic retries, fails until access resumes. Check the error body for that code and alert a human instead of retrying. A spend limit you set yourself returns HTTP 400 invalid_request_error, with a message beginning "You have reached your specified API usage limits".

Stop the broken server and restart the normal one with uvicorn main:app --reload. The Reference section has the full error table, the exception classes and a try/except pattern to wrap around your calls.

7. Set a spend limit and size your costs

Set a spend limit in the Console so a bug can't run up an open-ended bill, then check the numbers that drive your cost. The spend limit is the only lever that protects you from your own code, so do it first.

Open the Console, go to Settings > Billing, find the Spend limits section, and click Adjust limit (or Set limit if you haven't set one). Enter a value. Per the rate limits page, your spend limit can't exceed your current tier's cap, and the default Workspace can't have its own limits.

Checkpoint: the Spend limits section shows the value you entered, and Console > Usage lists the requests you made in steps 2 to 6.

Now the levers. Pick the model on purpose. This tutorial uses Sonnet 5.5 at $2 per million input tokens and $10 per million output, which the models overview describes as the best combination of speed and intelligence. In fairness, the same overview says to start with Claude Opus 5.5 for most workloads, so Anthropic's own default is the more expensive model. A triage classifier is a simple, high-volume task, so Sonnet is a defensible default and Haiku is worth testing. If you're coming from the previous Sonnet, what changed when Sonnet 5.5 replaced Sonnet 5 lists the differences.

Here's what one triage request costs, as plain arithmetic rather than a measured figure. 1,000 input tokens and 300 output tokens on Sonnet 5.5 is 1,000 x $2 / 1M plus 300 x $10 / 1M, which is $0.002 + $0.003 = $0.005 before thinking tokens. Adaptive thinking adds billed output tokens on top.

Treat max_tokens as a ceiling, and remember that thinking counts toward it. The thinking page says the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you, and they count toward max_tokens alongside the response text. Set it too low and a reply can end with stop_reason: "max_tokens" before the visible answer finishes, which step 4 turns into a 502. Prices for the other models, prompt caching and the batch discount are in Reference.

8. Run the whole feature and check a bad input

Start the finished server, run the triage and streaming flows end to end, then send a malformed request to confirm it's rejected. Your main.py and static/index.html now match the Python files in Full code.

From claude-quickstart/, with the virtual environment active:

Bash
uvicorn main:app --reload

In a second terminal, run the triage call:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "I was charged twice for my subscription this month."}'

You should see something like (the values will differ):

JSON
{"category":"billing","urgency":"medium","summary":"Customer reports being charged twice for this month's subscription."}

Open http://localhost:8000, paste the same sentence and click Summarise. You should see a two-sentence summary appear word by word.

Negative check: send a body with no text field.

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{}'

You should see something like (the exact fields depend on your Pydantic version):

JSON
{"detail":[{"type":"missing","loc":["body","text"],"msg":"Field required","input":{}}]}

FastAPI rejects it with a 422 before any tokens are spent.

Clean up

Stop uvicorn with Ctrl+C in its terminal and leave the virtual environment:

Bash
deactivate

To remove the project, delete the claude-quickstart folder. If you created the API key only for this tutorial, remove it from the Settings > API keys page where you made it. Leave the spend limit in place.

Full code

Steps 3 to 5 build one server in pieces. Here are the assembled files, ready to paste.

Python with FastAPI

The project layout:

text
claude-quickstart/
  main.py
  quickstart.py
  .env
  .env.example
  .gitignore
  static/
    index.html

Install and run, from the project folder with the virtual environment from step 1 active:

Bash
pip install anthropic fastapi uvicorn python-dotenv
cp .env.example .env   # then put your real key in .env
uvicorn main:app --reload

.env.example:

text
ANTHROPIC_API_KEY=

.gitignore:

text
.env
.venv/
__pycache__/

main.py:

Python
import json
from typing import Literal

import anthropic
from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from fastapi.staticfiles import StaticFiles
from pydantic import BaseModel, ValidationError

load_dotenv()

app = FastAPI()
client = anthropic.Anthropic()


class Feedback(BaseModel):
    text: str


class Triage(BaseModel):
    category: Literal["bug", "billing", "feature_request", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str


@app.post("/api/triage")
def triage(body: Feedback):
    try:
        response = client.messages.parse(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[{"role": "user", "content": "Triage this customer feedback. Be concise.\n\n" + body.text[:4000]}],
            output_format=Triage,
        )
    except ValidationError:
        raise HTTPException(status_code=502, detail="No usable result (output did not match the schema)")
    if response.stop_reason in ("max_tokens", "refusal") or response.parsed_output is None:
        raise HTTPException(status_code=502, detail=f"No usable result ({response.stop_reason}), request {response._request_id}")
    return response.parsed_output


@app.post("/api/triage/stream")
def triage_stream(body: Feedback):
    def generate():
        with client.messages.stream(
            model="claude-sonnet-5-5",
            max_tokens=1024,
            messages=[
                {
                    "role": "user",
                    "content": "Summarise this feedback in two sentences:\n\n" + body.text[:4000],
                }
            ],
        ) as stream:
            for text in stream.text_stream:
                yield f"data: {json.dumps(text)}\n\n"

    return StreamingResponse(generate(), media_type="text/event-stream")


app.mount("/", StaticFiles(directory="static", html=True), name="static")

static/index.html:

html
<!doctype html>
<html>
  <body>
    <textarea id="feedback" rows="6" cols="60"></textarea>
    <button id="go">Summarise</button>
    <div id="output" style="white-space: pre-wrap"></div>

    <script>
      const output = document.getElementById("output");

      document.getElementById("go").addEventListener("click", async () => {
        const text = document.getElementById("feedback").value;
        output.textContent = "";

        const res = await fetch("/api/triage/stream", {
          method: "POST",
          headers: { "content-type": "application/json" },
          body: JSON.stringify({ text })
        });

        const reader = res.body.getReader();
        const decoder = new TextDecoder();
        let buffer = "";

        while (true) {
          const { done, value } = await reader.read();
          if (done) break;
          buffer += decoder.decode(value, { stream: true });
          const events = buffer.split("\n\n");
          buffer = events.pop() ?? "";
          for (const event of events) {
            output.textContent += JSON.parse(event.slice("data: ".length));
          }
        }
      });
    </script>
  </body>
</html>

quickstart.py is the file from step 2. The server doesn't use it.

TypeScript with Express

This version is self-contained. In a new folder, create the project and install everything:

Bash
mkdir claude-quickstart-ts && cd claude-quickstart-ts
npm init -y
npm pkg set type=module
npm install @anthropic-ai/sdk express zod
npm install -D typescript tsx @types/express @types/node
mkdir static

The type=module line lets you use import and top-level await. Then create these files with the same contents as the Python version: .env.example, .env (with your real key), and static/index.html. For .gitignore, use:

text
.env
node_modules/

Run it:

Bash
npx tsx --env-file=.env server.ts

server.ts (Express 5, which is what npm install express installs):

TypeScript
import express from "express";
import Anthropic from "@anthropic-ai/sdk";
import { z } from "zod";
import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod";

const Triage = z.object({
  category: z.enum(["bug", "billing", "feature_request", "other"]),
  urgency: z.enum(["low", "medium", "high"]),
  summary: z.string()
});

const app = express();
app.use(express.json());

const client = new Anthropic();

app.post("/api/triage", async (req, res) => {
  const text = String(req.body?.text ?? "").slice(0, 4000);

  try {
    const response = await client.messages.parse({
      model: "claude-sonnet-5-5",
      max_tokens: 1024,
      messages: [{ role: "user", content: "Triage this customer feedback. Be concise.\n\n" + text }],
      output_config: { format: zodOutputFormat(Triage) }
    });

    if (response.stop_reason === "max_tokens" || response.stop_reason === "refusal" || !response.parsed_output) {
      res.status(502).json({ error: `No usable result (${response.stop_reason}), request ${response._request_id}` });
      return;
    }

    res.json(response.parsed_output);
  } catch (err) {
    console.error(err);
    res.status(502).json({ error: "Upstream error" });
  }
});

app.post("/api/triage/stream", async (req, res) => {
  const text = String(req.body?.text ?? "").slice(0, 4000);

  res.setHeader("content-type", "text/event-stream");

  try {
    await client.messages
      .stream({
        model: "claude-sonnet-5-5",
        max_tokens: 1024,
        messages: [{ role: "user", content: "Summarise this feedback in two sentences:\n\n" + text }]
      })
      .on("text", (chunk) => {
        res.write(`data: ${JSON.stringify(chunk)}\n\n`);
      })
      .finalMessage();
  } catch (err) {
    console.error(err);
  }

  res.end();
});

app.use(express.static("static"));

app.listen(Number(process.env.PORT ?? 3000));

To type-check without running the server:

Bash
npx tsc --noEmit --strict --module nodenext --target es2022 server.ts

Open http://localhost:3000. The page's fetch URL is relative, so the same index.html works against either server. The curl commands from the steps work with port 3000. If that port is taken, start the server with PORT=3001 npx tsx --env-file=.env server.ts and change the curl URL to match, the TypeScript counterpart of step 3's --port 8001 tip.

Reference

Error codes

HTTP codes and error types from the errors page:

HTTP codeError type
400invalid_request_error
401authentication_error
402billing_error
403permission_error
404not_found_error
409conflict_error
413request_too_large
429rate_limit_error
500api_error
504timeout_error
529overloaded_error

The body looks like {"type":"error","error":{"type":"not_found_error","message":"..."},"request_id":"req_..."}.

Exceptions, retries and timeouts

The Python exception classes map to status codes: 400 is BadRequestError, 401 AuthenticationError, 403 PermissionDeniedError, 404 NotFoundError, 409 ConflictError, 422 UnprocessableEntityError, 429 RateLimitError, and 500 and above InternalServerError. Connection failures raise APIConnectionError, and timeouts raise APITimeoutError in Python or APIConnectionTimeoutError in TypeScript. You change the retry count with max_retries in Python or maxRetries in TypeScript. Python's default timeout is 10 minutes.

For illustration only, not part of this project, the SDK page's pattern for catching typed errors in Python:

Python
import anthropic

client = anthropic.Anthropic(max_retries=2)

try:
    message = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello, Claude"}],
    )
except anthropic.APIConnectionError as e:
    print("The server could not be reached")
except anthropic.RateLimitError as e:
    print("A 429 status code was received; we should back off a bit.")
except anthropic.APIStatusError as e:
    print(e.status_code)
    print(e.response)

Two validation errors specific to Sonnet 5.5 are worth recognizing on sight. Sending thinking: {"type": "disabled"} returns a 400, and the message tells you to send "thinking": {"type": "between_tools"} to turn thinking off on this model. Non-default temperature, top_p or top_k also returns a 400.

Rate limits

The rate limits page measures limits in requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM), with capacity continuously replenished up to your maximum, like a token bucket. The response headers tell you where you stand: retry-after, anthropic-ratelimit-requests-remaining, anthropic-ratelimit-input-tokens-remaining, anthropic-ratelimit-output-tokens-remaining and anthropic-ratelimit-requests-reset.

New organizations may start in the Evaluation tier, with limits below the standard ones. For scale, the page listed Start-tier limits for Sonnet 5.5 of 1,000 RPM, 2,000,000 ITPM and 400,000 OTPM when checked on 2026-10-01. Spend limits can't exceed your tier's cap: $500 a month on Start, $1,000 on Build and $200,000 on Scale. For per-project limits, use Workspaces, though you can't set limits on the default Workspace.

Structured output limits

The structured outputs page lists several limits:

  • Numerical constraints (minimum, maximum, multipleOf) aren't supported.
  • additionalProperties must be set to false for objects.
  • If Claude refuses a request, you get stop_reason: "refusal" and the output may not match your schema.
  • If the response is cut off at max_tokens, the output may be incomplete.
  • Claude may return enum values that differ only in capitalization, so compare enum values case-insensitively.

The SDK transforms your schema, sends it as output_config.format, validates the response and hands back the parsed model in parsed_output. If you call the HTTP API directly, output_config.format is the parameter you set yourself. The docs headline the feature as guaranteed, schema-validated JSON. Treat the last three limits as the fine print on that guarantee.

Model prices

The pricing page lists these prices per million tokens (MTok), checked 2026-10-01:

ModelModel IDInputOutput
Sonnet 5.5claude-sonnet-5-5$2$10
Opus 5.5claude-opus-5-5$4$20
Haiku 4.5claude-haiku-4-5-20251001$1$5
Fable 5.1claude-fable-5-1$10$50

The page's short version: choose Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning. Haiku 4.5 is listed as retiring no sooner than 15 October 2026, so check the deprecations page before choosing it. The same 1,000 input and 300 output token request costs $0.001 + $0.0015 = $0.0025 on Haiku 4.5.

The page's rule of thumb for sizing text is that 1 token is approximately 4 characters or 0.75 words in English, but that's optimistic for newer tokenizers: the page notes about 30% more tokens on Claude 4.7+ tokenizers. You can count tokens before you send. In Python, client.messages.count_tokens takes the same model and messages and returns the input size in count.input_tokens.

Sonnet 5.5's default effort is high, and output_config.effort is the control for it. Setting thinking to {"type": "between_tools"} is the lowest setting, and it works at high effort or below. Treat both as options to test, since the docs don't quantify the savings.

Prompt caching and batches

The prompt caching page covers the cache_control parameter, set to {"type": "ephemeral"} in Python or { type: "ephemeral" } in TypeScript. The default lifetime is 5 minutes, and you can ask for an hour with { "cache_control": { "type": "ephemeral", "ttl": "1h" } }. On Sonnet 5.5, a 5-minute cache write costs $2.50 per MTok, a 1-hour write $4, and a cache read $0.20. The page says caching pays off after one cache read for the 5-minute duration, thanks to the 1.25x write price. The usage object reports cache_creation_input_tokens, cache_read_input_tokens and input_tokens so you can verify it's working.

The minimum cacheable prompt is 512 tokens on Sonnet 5.5 and 4,096 on Haiku 4.5. The one-line prompts in this tutorial are far below that, so caching does nothing for them. It earns its keep when a long fixed prefix, such as a style guide or a few dozen examples, repeats across many requests. Per the rate limits page, cache reads also don't count toward ITPM for most models.

Batches suit work nobody's waiting on. The pricing page lists a 50% discount on both input and output tokens for asynchronous batch work, which puts Sonnet 5.5 at $1 and $5 per MTok.

Streaming helpers

When you need the complete message at the end of a stream, for logging token usage, say, Python has stream.get_final_message() and TypeScript has await stream.finalMessage().

Ship checklist

Before this goes to real users, check every line below.

  • The key lives only in server environment variables or a secrets manager, never in frontend code or source control.
  • Each environment (local, staging, production) has its own key, and deployed code runs on a service account key.
  • .env is in .gitignore.
  • Your endpoint authenticates the caller, caps input length and rate limits per user.
  • Every request sets max_tokens, and it's large enough to cover thinking plus the answer.
  • Your code loops over message.content and checks block.type rather than reading index zero.
  • You handle stop_reason values max_tokens and refusal.
  • You catch RateLimitError, read retry-after, and check the error body for enforced_spend_limit_reached and alert a human instead of retrying.
  • You log _request_id on every failure.
  • A spend limit is set under Settings > Billing.
  • You pin claude-sonnet-5-5. The models overview says every Claude model ID is a pinned snapshot, so behavior won't shift under you, and the model deprecations page lists Sonnet 5.5's retirement as not sooner than September 28, 2027. Check it periodically, and re-run your tests whenever you change the model ID.

Common mistakes

  • Reading message.content[0].text. With adaptive thinking on, the first block can be a thinking block, so loop over the blocks and check type (step 2).
  • Putting the key in frontend code. Keep it on the server and call your own endpoint (step 3).
  • Setting temperature, top_p or top_k to a non-default value. On Sonnet 5.5 that returns a 400 (step 2).
  • Retrying a spend-cap 429. It has no retry-after header and fails until access resumes, so alert a person instead (step 6).
  • Setting max_tokens too low. Thinking tokens count toward it, so the reply can stop at max_tokens before the visible answer is done (steps 4 and 7).

Next steps

  • Wrap your Claude calls in the try/except pattern from Reference, so a failed call returns your own error instead of a 500.
  • Add tool use, so the model can call your own functions. The Sonnet 5.5 overview says text between tool calls returns in thinking blocks, so the loop-over-blocks habit from step 2 matters even more there.
  • Add conversation history. The Messages API is stateless, so a follow-up question means appending each turn to the messages list and sending all of it again.
  • Backfill old feedback in bulk. Batch work gets the 50% discount in Reference, which puts Sonnet 5.5 at $1 and $5 per MTok.

FAQ

Is the Claude API free?

No, it's pay-per-token. New users receive a small amount of free credits to test the API, according to Anthropic's pricing FAQ. After that you're billed on input and output tokens. Set a spend limit under Settings > Billing so a bug can't run up an open-ended bill.

Can I call the Claude API directly from the browser?

Not by default. The TypeScript SDK disables browser use to avoid exposing your secret API credentials. The dangerouslyAllowBrowser option exists, but the SDK documentation warns that any user with access to the browser can inspect, extract and misuse the credentials. Route calls through your own backend, as in step 3.

Which Claude model should a beginner start with?

This tutorial uses claude-sonnet-5-5, at $2 per million input tokens and $10 per million output, with a 1M context window and 128K max output. Anthropic's docs lead with Opus 5.5 as the default starting point, at $4 and $20. Start with Sonnet to keep costs low, move to Opus if quality falls short, and test Haiku for simple, high-volume tasks.

Do I need the SDK to use the Claude API?

No. You can send a POST to https://api.anthropic.com/v1/messages with the headers x-api-key, anthropic-version: 2023-06-01 and content-type: application/json. The SDK adds typed responses, automatic retries and streaming helpers, and sends the version header for you, so most projects are better off using it.

Why does message.content[0].text fail on Sonnet 5.5?

Adaptive thinking is on by default on Sonnet 5.5, and the thinking arrives in thinking content blocks ahead of the response. Index zero can therefore be a thinking block with no text field. Loop over message.content and handle only blocks whose type is "text", as in step 2.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now