October 1, 2026

How to use the OpenAI Responses API: build an AI feature step by step

Build an AI feature on the OpenAI Responses API: first request, streaming, structured output, function calling, conversation state and cost limits, with code.

Guide

Tech

This is how to use the OpenAI Responses API in a real project, not a throwaway script. You'll end up with a support-ticket triage assistant that returns JSON your code can trust, streams a summary into the browser as it's written, looks up an order through a function call, answers follow-up questions without resending the conversation, and keeps the API key on your server. If you're deciding who on a team gets keys and under what rules, AI in the engineering workflow: tools, permissions and policy covers that side.

What you'll build

A triage assistant for one ticket: "The invoice page crashes when I click download." Your backend takes the ticket text, returns {category, urgency, summary}, streams a short summary to a small web page, fetches order details through a function the model can call, and handles a follow-up question about the customer. Plan on about an hour. The code is Python with FastAPI. A TypeScript version is cut from this article for length, and the SDK steps map one to one if you use Node.

Here's the finished endpoint, called with curl. The values in the response are an example of the shape, not real output:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'
JSON
{
  "category": "bug",
  "urgency": "high",
  "summary": "Customer reports the invoice page crashes when they click download."
}

What you need before you start

  • An OpenAI platform account with billing set up (you'll create the key in step 2).
  • Python 3.10 or higher, which is the Python SDK's stated requirement. This tutorial uses openai 3.22.1, the version PyPI listed on 2026-10-01, and pins the major version.
  • A macOS or Linux shell. On Windows, use WSL, because the commands use source, cp and export.
  • A terminal and a project folder you don't mind filling with experiments.

Quickstart

Already have a key? Install the SDK, export the key, save this as quickstart.py and run python quickstart.py:

Bash
pip install "openai>=3.22,<4"
export OPENAI_API_KEY="your-api-key-here"
Python
from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    input="Summarise this support ticket in one sentence: The invoice page crashes when I click download.",
)

print(response.output_text)

Three things to get right before you build on it:

  • Keep the key on the server. Never put it in frontend code or source control.
  • Set a hard spend limit, so a bug can't run up an open-ended bill. A spend alert only notifies you.
  • Check response.status and the items in response.output, not just output_text. A reply can be cut off or refused and still arrive with HTTP 200.

How to use the Responses API in your own project

1. Create the project and install the packages

You need a folder, a virtual environment, every package the article uses, a .env file for the key and a .gitignore that keeps that file out of git. Run these in your terminal, from wherever you keep projects:

Bash
mkdir -p triage-assistant/static && cd triage-assistant
python3 -m venv .venv && source .venv/bin/activate
pip install "openai>=3.22,<4" fastapi uvicorn python-dotenv
printf 'OPENAI_API_KEY=\n' > .env.example
cp .env.example .env
printf '.env\n.venv/\n__pycache__/\n' > .gitignore

The <4 cap keeps a future major release from changing your imports under you. Bump it on purpose, after reading the changelog. .env is in .gitignore before your first commit, which matters: a key in git history is a leaked key, even if you delete the file in the next commit.

Check that the packages import:

Bash
python -c "import openai, fastapi, uvicorn, dotenv; print('imports ok')"

You should see something like:

text
imports ok

If you see ModuleNotFoundError, the virtual environment isn't active. Run source .venv/bin/activate again from triage-assistant/.

2. Create an API key and a spend limit, then load it

You need one credential and one safety net. Open the OpenAI dashboard and create an API key, following the quickstart, and copy it straight away. Then open .env and replace the empty value:

text
OPENAI_API_KEY=your-api-key-here

The limits live in the dashboard's limits section under your organization settings. The rate limits guide explains them, and menu wording changes, so follow it rather than a screenshot from an older tutorial. There are two kinds. A spend alert only notifies you. A hard limit starts returning 429 errors once you hit the threshold. Set the hard limit now, before more code exists, because it's the only control that protects you from your own bugs.

The SDK reads OPENAI_API_KEY from the shell environment, not from the .env file, so the code calls load_dotenv() before it creates the client. The Python SDK README suggests python-dotenv for this. Check that the key loads, from triage-assistant/:

Bash
python -c "import os; from dotenv import load_dotenv; load_dotenv(); print(bool(os.environ.get('OPENAI_API_KEY')))"

You should see:

text
True

If you see False, the OPENAI_API_KEY= line in .env is still empty, or you're not in the project folder.

3. Send your first Responses request

Create a client, call responses.create with a model ID, input and instructions, then read output_text. If it prints a sentence, your key, install and network path all work.

Create quickstart.py in triage-assistant/:

Python
from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    instructions="You triage customer support tickets. Be concise.",
    input="The invoice page crashes when I click download.",
)

print(response.status)
print(response.output_text)

Run it:

Bash
python quickstart.py

You should see something like this (example output, the wording will differ):

text
completed
The customer can't download invoices because the invoice page crashes. Treat it as a high-priority bug.

If you see AuthenticationError, the key isn't loaded or isn't valid. Redo step 2.

The model ID is gpt-6.1-sol. OpenAI's models page recommends gpt-6-astra as the default and also lists gpt-6.1-sol and gpt-6-luna, all with a 1.05M-token context window. This tutorial picks Sol because the docs position it as the balance of intelligence and cost. If ticket classification turns out to need more reasoning, swap the ID in one place. The Reference section compares the options.

The reply is a response object with an id, a status, an output list of items, and the convenience property output_text. Why does output_text come with a warning? It's a shortcut over the output items, and a response can contain items that aren't text, such as a refusal or a function call. You'll handle those in steps 5 and 7.

4. Put the call behind your own endpoint

Your browser should call an endpoint you control, and only that endpoint should talk to OpenAI. That keeps the key on the server and gives you one place for auth, input limits and rate limiting.

text
browser
  |  POST /api/triage (no key in the page)
  v
your endpoint: auth, input length cap, per-user rate limit
  |  OPENAI_API_KEY (lives only on the server)
  v
OpenAI Responses API  (POST /v1/responses)

The Node SDK makes the point for you. Its README says the dangerouslyAllowBrowser option "can be dangerous because it exposes your secret API credentials." A key in frontend code is a key you've published. If you want the longer version of how that failure shows up in shipped apps, what to check before a vibe coded app reaches real users covers it.

The FastAPI routes, the order lookup stub, the server-sent event framing and the browser fetch code in steps 4 to 9 are tutorial code, not OpenAI SDK calls. They're minimal, so read them as a starting point.

Create main.py in triage-assistant/. One rule holds for the whole article: app.mount(...) stays the last line of the file, and every route you add later goes above it. A mount at / matches everything, so a route placed below it never runs and answers 405.

Python
from dotenv import load_dotenv
from fastapi import FastAPI
from fastapi.staticfiles import StaticFiles
from openai import OpenAI
from pydantic import BaseModel

load_dotenv()

MODEL = "gpt-6.1-sol"
INSTRUCTIONS = "You triage customer support tickets. Be concise."

app = FastAPI()
client = OpenAI()


class Ticket(BaseModel):
    text: str


@app.post("/api/triage")
def triage(body: Ticket):
    response = client.responses.create(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.text[:4000],
    )
    return {"text": response.output_text}


app.mount("/", StaticFiles(directory="static", html=True), name="static")

Start the server in a terminal and leave it running. --reload restarts it when you save a file:

Bash
uvicorn main:app --reload

In a second terminal, call the endpoint with the curl command from the top of the article:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (example output):

text
{"text":"Bug: the invoice page crashes on download. Likely high urgency."}

If you see Address already in use, another process holds port 8000. Stop it, or start uvicorn with --port 8002 and change the curl URL.

The body.text[:4000] slice is the input length cap. Without one, anyone who finds your endpoint can send a 10 MB document and you pay for it. This endpoint has no authentication, which is fine on localhost and not fine anywhere else. Put your session or token check in front of it, and add a per-user rate limit, before you deploy.

5. Return structured output

Ask the SDK for a typed result instead of parsing free text. You define the shape once as a Pydantic model, call responses.parse with text_format, and read the validated object from output_parsed. The structured outputs guide documents the call. On the wire, text_format becomes the text.format parameter, which is where response_format went in the move from Chat Completions.

The guide names two ways a structured request can fail without an HTTP error. Check response.status == "incomplete", where incomplete_details.reason can be max_output_tokens, and look for a content item with type == "refusal". Either way, the output may not match your schema. The check helper below, which is tutorial code, turns both into a 502.

Replace the whole of main.py with this version. It adds the Triage model and check, and swaps the triage function to responses.parse. The mount is still last:

Python
from typing import Literal

from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.staticfiles import StaticFiles
from openai import OpenAI
from pydantic import BaseModel

load_dotenv()

MODEL = "gpt-6.1-sol"
INSTRUCTIONS = "You triage customer support tickets. Be concise."

app = FastAPI()
client = OpenAI()


class Ticket(BaseModel):
    text: str


class Triage(BaseModel):
    category: Literal["bug", "billing", "feature_request", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str


def check(response):
    if response.status == "incomplete":
        reason = response.incomplete_details.reason if response.incomplete_details else None
        raise HTTPException(502, f"Incomplete ({reason}), request {response._request_id}")
    for item in response.output:
        if item.type == "message":
            for part in item.content:
                if part.type == "refusal":
                    raise HTTPException(502, f"Refused, request {response._request_id}")


@app.post("/api/triage")
def triage(body: Ticket):
    response = client.responses.parse(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.text[:4000],
        text_format=Triage,
        max_output_tokens=2000,
    )
    check(response)
    if response.output_parsed is None:
        raise HTTPException(502, f"No parsed result, request {response._request_id}")
    return response.output_parsed


app.mount("/", StaticFiles(directory="static", html=True), name="static")

max_output_tokens caps what the model can produce. Leave generous headroom: a cap that's too tight is exactly what produces the incomplete status. check gets reused in steps 7 and 8.

Run the same curl as step 4 against the reloaded server:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (example output, the values will vary):

text
{"category":"bug","urgency":"high","summary":"Customer reports the invoice page crashes when they click download."}

If you see Internal Server Error, read the uvicorn terminal. An AuthenticationError there means the key is wrong (redo step 2). Otherwise a missing import is the usual cause, so check the file against the block above.

6. Stream the summary to the browser

Streaming sends text to the user as the model writes it, so a longer answer starts appearing early. Set stream=True and read the events. The streaming guide lists the ones you need: response.created, response.output_text.delta, response.completed and error.

This takes three edits to main.py and one new file.

First, replace the import block at the top of main.py with this one. It adds json, StreamingResponse and APIError:

Python
import json
from typing import Literal

from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from fastapi.staticfiles import StaticFiles
from openai import APIError, OpenAI
from pydantic import BaseModel

Second, paste this helper below the check function and above the first @app.post line:

Python
def sse(payload: dict) -> str:
    return f"data: {json.dumps(payload)}\n\n"

Third, paste this route directly above the app.mount(...) line. It passes each text delta to the browser as a server-sent event:

Python
@app.post("/api/triage/stream")
def triage_stream(body: Ticket):
    def generate():
        try:
            stream = client.responses.create(
                model=MODEL,
                instructions="Summarise this support ticket in two sentences.",
                input=body.text[:4000],
                stream=True,
                max_output_tokens=2000,
            )
            for event in stream:
                if event.type == "response.output_text.delta":
                    yield sse({"delta": event.delta})
                elif event.type in ("error", "response.failed"):
                    yield sse({"error": "The stream failed."})
                    return
                elif event.type == "response.completed":
                    yield sse({"done": True})
        except APIError:
            yield sse({"error": "The stream failed."})

    return StreamingResponse(generate(), media_type="text/event-stream")

Each frame is JSON, so newlines inside the text can't break the event framing. The done frame matters: a stream can start fine and still fail partway, so the UI needs a state for "this stopped early", not only success and failure.

Now the page. A plain fetch with a stream reader does the job (EventSource only sends GET requests, and you're posting a body). Create static/index.html:

html
<!doctype html>
<html>
  <body>
    <textarea id="ticket" rows="6" cols="60">The invoice page crashes when I click download.</textarea>
    <button id="go">Summarise</button>
    <div id="output" style="white-space: pre-wrap"></div>

    <script>
      const output = document.getElementById("output");

      document.getElementById("go").addEventListener("click", async () => {
        const text = document.getElementById("ticket").value;
        output.textContent = "";
        let finished = false;

        const res = await fetch("/api/triage/stream", {
          method: "POST",
          headers: { "content-type": "application/json" },
          body: JSON.stringify({ text })
        });
        if (!res.ok) { output.textContent = "Request failed (" + res.status + ")."; return; }

        const reader = res.body.getReader();
        const decoder = new TextDecoder();
        let buffer = "";

        while (true) {
          const { done, value } = await reader.read();
          if (done) break;
          buffer += decoder.decode(value, { stream: true });
          const events = buffer.split("\n\n");
          buffer = events.pop() ?? "";
          for (const event of events) {
            const data = JSON.parse(event.slice("data: ".length));
            if (data.delta) output.textContent += data.delta;
            if (data.error) output.textContent += "\n[" + data.error + "]";
            if (data.done) finished = true;
          }
        }

        if (!finished) output.textContent += "\n[Stopped early]";
      });
    </script>
  </body>
</html>

Check the stream from the terminal first. The -N flag turns off curl's buffering so frames print as they arrive:

Bash
curl -N -X POST http://localhost:8000/api/triage/stream \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (example output, the text will differ):

text
data: {"delta": "The invoice"}

data: {"delta": " page crashes when a customer clicks download."}

data: {"done": true}

Then open http://localhost:8000 and click Summarise. The text should appear in pieces. The relative fetch URL works only on the same origin, or through a dev-server proxy. Don't "fix" a cross-origin error by opening the API to every origin with allow_origins=["*"].

If the page shows Request failed (404), the route is missing. If it shows 405, the new route sits below the app.mount(...) line. Move it above.

7. Add the order lookup tool

Function calling lets the model ask your code for data it doesn't have. Here the model gets one tool, lookup_order(order_id), so triage can fetch the customer's order. You run the function, send the result back, and the model finishes its answer.

The function calling guide defines the shapes. A tool is {"type": "function", "name": ..., "description": ..., "parameters": ...}. When the model wants one, the output contains an item of type function_call. You answer with an input item {"type": "function_call_output", "call_id": ..., "output": ...}, using the call_id from the call. Setting strict: true requires additionalProperties: false with every field listed as required.

The latest-model guide says to use the Responses API for tool calling, and that Chat Completions supports requests without tools on gpt-6-astra and gpt-6.1-sol. If you're moving existing tool-calling code across, our GPT-6 Astra release breakdown has the migration checklist, so this tutorial doesn't repeat it.

Three pastes into main.py, all tutorial code. The stub is a hard-coded dictionary standing in for your order system, and the loop carries a round cap, because a model that keeps calling tools would otherwise keep you paying.

First, paste this block below the INSTRUCTIONS = ... line and above app = FastAPI():

Python
MAX_TOOL_ROUNDS = 4

ORDERS = {
    "A-1001": {"customer": "Dana Lee", "plan": "Pro", "status": "paid"},
}

TOOLS = [
    {
        "type": "function",
        "name": "lookup_order",
        "description": "Fetch an order by its ID, including the customer's plan.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
            "additionalProperties": False,
        },
        "strict": True,
    }
]

Second, paste this function below the Triage class and above def check:

Python
def lookup_order(order_id: str) -> dict:
    return ORDERS.get(order_id, {"error": "order not found"})

Third, paste this route directly above the app.mount(...) line, below the stream route:

Python
@app.post("/api/triage/order")
def triage_with_order(body: Ticket):
    response = client.responses.create(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.text[:4000],
        tools=TOOLS,
        max_output_tokens=2000,
    )
    for _ in range(MAX_TOOL_ROUNDS):
        calls = [item for item in response.output if item.type == "function_call"]
        if not calls:
            check(response)
            return {"response_id": response.id, "text": response.output_text}
        results = []
        for call in calls:
            result = lookup_order(**json.loads(call.arguments))
            results.append(
                {
                    "type": "function_call_output",
                    "call_id": call.call_id,
                    "output": json.dumps(result),
                }
            )
        response = client.responses.create(
            model=MODEL,
            instructions=INSTRUCTIONS,
            input=results,
            tools=TOOLS,
            previous_response_id=response.id,
            max_output_tokens=2000,
        )
    raise HTTPException(502, "Too many tool rounds")

The second call passes only the tool results, plus previous_response_id, which is the subject of step 8. Chaining means the server already holds the model's earlier items. Treat the arguments the model passes to lookup_order as untrusted input, and validate them against your real order service before you swap the stub out.

Try it with a ticket that names an order:

Bash
curl -X POST http://localhost:8000/api/triage/order \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download. Order A-1001."}'

You should see something like (example output, copy the real response_id from yours):

text
{"response_id":"<your response id>","text":"Bug, high urgency. Order A-1001 belongs to Dana Lee on the Pro plan, so the crash blocks a paying customer."}

The model decides whether to call the tool, so a reply may not mention the plan every time. If it doesn't, send the request again.

8. Add follow-up questions with previous_response_id

The Responses API can hold the conversation for you, so a follow-up question doesn't mean resending the ticket. Pass previous_response_id=response.id on the next call. This tutorial uses that option, because the triage follow-up is short. The Reference section compares it with the Conversations API and manual replay.

Two things to know before you rely on it. It doesn't carry over top-level instructions, so the route resends them. And every prior input token in the chain is billed as input again, so a chain ten turns long re-bills the earlier turns on every call. Long chats get expensive in a way a single request doesn't suggest.

Two pastes into main.py. First, paste this class between the Ticket class and the Triage class:

Python
class FollowUp(BaseModel):
    response_id: str
    question: str

Second, paste this route directly above the app.mount(...) line, below the order route:

Python
@app.post("/api/triage/followup")
def followup(body: FollowUp):
    response = client.responses.create(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.question[:1000],
        previous_response_id=body.response_id,
        max_output_tokens=2000,
    )
    check(response)
    return {"response_id": response.id, "text": response.output_text}

Take the response_id from the step 7 reply and ask the question. Replace PASTE_ID_FROM_STEP_7 with the real id:

Bash
curl -X POST http://localhost:8000/api/triage/followup \
  -H "content-type: application/json" \
  -d '{"response_id": "PASTE_ID_FROM_STEP_7", "question": "Which plan is that customer on?"}'

You should see something like (example output):

text
{"response_id":"<a new response id>","text":"Dana Lee is on the Pro plan."}

The model answers from the earlier tool result, without the ticket being sent again. If you get Internal Server Error, you most likely sent the placeholder text instead of a real id.

9. Handle OpenAI errors

Catch the SDK's typed exceptions so a failed call becomes a clean response instead of a stack trace, and log the request ID. The client already retries twice by default on transient failures, so these handlers only see what retries couldn't fix.

First, replace the import block at the top of main.py with the final one. It adds JSONResponse and RateLimitError:

Python
import json
from typing import Literal

from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.responses import JSONResponse, StreamingResponse
from fastapi.staticfiles import StaticFiles
from openai import APIError, OpenAI, RateLimitError
from pydantic import BaseModel

Second, paste these handlers below the sse function and above the first @app.post line. This is tutorial code:

Python
@app.exception_handler(RateLimitError)
def on_rate_limit(request, exc):
    return JSONResponse(status_code=503, content={"error": "Busy, try again shortly."})


@app.exception_handler(APIError)
def on_api_error(request, exc):
    print(f"OpenAI error: {exc!r}")
    return JSONResponse(status_code=502, content={"error": "Upstream error."})

Six different things produce a 429, and only some of them clear if you wait. A spend-limit or credit 429 won't clear in the seconds a retry takes, and the SDK's automatic retries just burn time on it, so alert a human instead. The Reference section lists each kind.

To see a handler fire, start a second copy of the server with a deliberately wrong key. A key set inline overrides .env, because load_dotenv() doesn't replace a variable that's already set. In a second terminal, from triage-assistant/:

Bash
OPENAI_API_KEY=invalid-key uvicorn main:app --port 8001

In a third terminal:

Bash
curl -X POST http://localhost:8001/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download."}'

You should see something like (example output):

text
{"error":"Upstream error."}

The server terminal prints an OpenAI error: line with the exception. Stop the port 8001 server with Ctrl+C. If you see a stack trace instead of the JSON, the APIError handler isn't registered, so check the paste position.

10. Run and verify the whole project

Confirm the finished project end to end, with the real key and the server from step 4 still running. If you stopped it, start it again from triage-assistant/ with the virtual environment active:

Bash
uvicorn main:app --reload

Open http://localhost:8000, keep the sample ticket, and click Summarise. The summary should stream in piece by piece and end without a [Stopped early] line. Then run the three JSON routes in a second terminal, in this order:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download. Order A-1001."}'

curl -X POST http://localhost:8000/api/triage/order \
  -H "content-type: application/json" \
  -d '{"text": "The invoice page crashes when I click download. Order A-1001."}'

curl -X POST http://localhost:8000/api/triage/followup \
  -H "content-type: application/json" \
  -d '{"response_id": "PASTE_ID_FROM_ORDER_REPLY", "question": "Which plan is that customer on?"}'

You should see something like (example output, three replies in order):

text
{"category":"bug","urgency":"high","summary":"Customer reports the invoice page crashes when they click download."}
{"response_id":"<your response id>","text":"Bug, high urgency. Order A-1001 belongs to Dana Lee on the Pro plan."}
{"response_id":"<a new response id>","text":"Dana Lee is on the Pro plan."}

Now the negative check. Send a body with no text field, which your endpoint should reject before it spends a token:

Bash
curl -X POST http://localhost:8000/api/triage \
  -H "content-type: application/json" \
  -d '{}'

FastAPI's request validation answers with HTTP 422 and a JSON body that names the missing text field. The request never reaches OpenAI. Your main.py now matches the Full code section below.

Clean up

Stop the server with Ctrl+C in its terminal, then leave the virtual environment:

Bash
deactivate

To remove the project, run rm -rf triage-assistant from its parent folder. Delete the API key in the OpenAI dashboard if it was only for this tutorial. The spend limit can stay.

Full code

Steps 4 to 9 build one server in pieces, and main.py below is the assembled result. If your file differs from it, replace yours with this version. It's Python only.

The project layout:

text
triage-assistant/
  main.py
  quickstart.py
  .env
  .env.example
  .gitignore
  static/
    index.html

Install and run, from a macOS or Linux shell:

Bash
mkdir -p triage-assistant/static && cd triage-assistant
python3 -m venv .venv && source .venv/bin/activate
pip install "openai>=3.22,<4" fastapi uvicorn python-dotenv
cp .env.example .env   # then put your real key in .env
uvicorn main:app --reload

.env.example:

text
OPENAI_API_KEY=

.gitignore:

text
.env
.venv/
__pycache__/

quickstart.py:

Python
from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    instructions="You triage customer support tickets. Be concise.",
    input="The invoice page crashes when I click download.",
)

print(response.status)
print(response.output_text)

main.py:

Python
import json
from typing import Literal

from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from fastapi.responses import JSONResponse, StreamingResponse
from fastapi.staticfiles import StaticFiles
from openai import APIError, OpenAI, RateLimitError
from pydantic import BaseModel

load_dotenv()

MODEL = "gpt-6.1-sol"
INSTRUCTIONS = "You triage customer support tickets. Be concise."
MAX_TOOL_ROUNDS = 4

ORDERS = {
    "A-1001": {"customer": "Dana Lee", "plan": "Pro", "status": "paid"},
}

TOOLS = [
    {
        "type": "function",
        "name": "lookup_order",
        "description": "Fetch an order by its ID, including the customer's plan.",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
            "additionalProperties": False,
        },
        "strict": True,
    }
]

app = FastAPI()
client = OpenAI()


class Ticket(BaseModel):
    text: str


class FollowUp(BaseModel):
    response_id: str
    question: str


class Triage(BaseModel):
    category: Literal["bug", "billing", "feature_request", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str


def lookup_order(order_id: str) -> dict:
    return ORDERS.get(order_id, {"error": "order not found"})


def check(response):
    if response.status == "incomplete":
        reason = response.incomplete_details.reason if response.incomplete_details else None
        raise HTTPException(502, f"Incomplete ({reason}), request {response._request_id}")
    for item in response.output:
        if item.type == "message":
            for part in item.content:
                if part.type == "refusal":
                    raise HTTPException(502, f"Refused, request {response._request_id}")


def sse(payload: dict) -> str:
    return f"data: {json.dumps(payload)}\n\n"


@app.exception_handler(RateLimitError)
def on_rate_limit(request, exc):
    return JSONResponse(status_code=503, content={"error": "Busy, try again shortly."})


@app.exception_handler(APIError)
def on_api_error(request, exc):
    print(f"OpenAI error: {exc!r}")
    return JSONResponse(status_code=502, content={"error": "Upstream error."})


@app.post("/api/triage")
def triage(body: Ticket):
    response = client.responses.parse(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.text[:4000],
        text_format=Triage,
        max_output_tokens=2000,
    )
    check(response)
    if response.output_parsed is None:
        raise HTTPException(502, f"No parsed result, request {response._request_id}")
    return response.output_parsed


@app.post("/api/triage/stream")
def triage_stream(body: Ticket):
    def generate():
        try:
            stream = client.responses.create(
                model=MODEL,
                instructions="Summarise this support ticket in two sentences.",
                input=body.text[:4000],
                stream=True,
                max_output_tokens=2000,
            )
            for event in stream:
                if event.type == "response.output_text.delta":
                    yield sse({"delta": event.delta})
                elif event.type in ("error", "response.failed"):
                    yield sse({"error": "The stream failed."})
                    return
                elif event.type == "response.completed":
                    yield sse({"done": True})
        except APIError:
            yield sse({"error": "The stream failed."})

    return StreamingResponse(generate(), media_type="text/event-stream")


@app.post("/api/triage/order")
def triage_with_order(body: Ticket):
    response = client.responses.create(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.text[:4000],
        tools=TOOLS,
        max_output_tokens=2000,
    )
    for _ in range(MAX_TOOL_ROUNDS):
        calls = [item for item in response.output if item.type == "function_call"]
        if not calls:
            check(response)
            return {"response_id": response.id, "text": response.output_text}
        results = []
        for call in calls:
            result = lookup_order(**json.loads(call.arguments))
            results.append(
                {
                    "type": "function_call_output",
                    "call_id": call.call_id,
                    "output": json.dumps(result),
                }
            )
        response = client.responses.create(
            model=MODEL,
            instructions=INSTRUCTIONS,
            input=results,
            tools=TOOLS,
            previous_response_id=response.id,
            max_output_tokens=2000,
        )
    raise HTTPException(502, "Too many tool rounds")


@app.post("/api/triage/followup")
def followup(body: FollowUp):
    response = client.responses.create(
        model=MODEL,
        instructions=INSTRUCTIONS,
        input=body.question[:1000],
        previous_response_id=body.response_id,
        max_output_tokens=2000,
    )
    check(response)
    return {"response_id": response.id, "text": response.output_text}


app.mount("/", StaticFiles(directory="static", html=True), name="static")

static/index.html:

html
<!doctype html>
<html>
  <body>
    <textarea id="ticket" rows="6" cols="60">The invoice page crashes when I click download.</textarea>
    <button id="go">Summarise</button>
    <div id="output" style="white-space: pre-wrap"></div>

    <script>
      const output = document.getElementById("output");

      document.getElementById("go").addEventListener("click", async () => {
        const text = document.getElementById("ticket").value;
        output.textContent = "";
        let finished = false;

        const res = await fetch("/api/triage/stream", {
          method: "POST",
          headers: { "content-type": "application/json" },
          body: JSON.stringify({ text })
        });
        if (!res.ok) { output.textContent = "Request failed (" + res.status + ")."; return; }

        const reader = res.body.getReader();
        const decoder = new TextDecoder();
        let buffer = "";

        while (true) {
          const { done, value } = await reader.read();
          if (done) break;
          buffer += decoder.decode(value, { stream: true });
          const events = buffer.split("\n\n");
          buffer = events.pop() ?? "";
          for (const event of events) {
            const data = JSON.parse(event.slice("data: ".length));
            if (data.delta) output.textContent += data.delta;
            if (data.error) output.textContent += "\n[" + data.error + "]";
            if (data.done) finished = true;
          }
        }

        if (!finished) output.textContent += "\n[Stopped early]";
      });
    </script>
  </body>
</html>

Open http://localhost:8000 for the streaming page. The JSON routes are POST /api/triage, /api/triage/order and /api/triage/followup, as in the curl examples above.

Reference

Error codes and exception classes

The error codes guide lists what comes back:

HTTP codeMeaning
401Invalid authentication
403Unsupported region
429Credit exhausted, rate limit, slow down, organization spend limit, project spend limit, or usage limit
500Server error
503Overloaded

The Python exception classes are APIConnectionError, APITimeoutError, AuthenticationError, BadRequestError, ConflictError, InternalServerError, NotFoundError, PermissionDeniedError, RateLimitError and UnprocessableEntityError.

Retries, timeouts and the kinds of 429

Per the Python SDK README and the Node SDK README, the client retries twice by default with exponential backoff, on connection errors, 408, 409, 429 and 500 and above. You change the count with max_retries in Python or maxRetries in Node, for example OpenAI(max_retries=2). Two is already the default, so only change it with a reason. The default timeout is 10 minutes. On any response, response._request_id carries the x-request-id header, so log it with every failure.

A rate-limit or slow-down 429 clears when you wait and follow Retry-After. Credit exhaustion needs you to add credits. A spend-limit 429 lasts until you raise the limit or the monthly limit resets. A usage-limit 429 needs a higher approved limit from OpenAI. Retrying the last three, including the SDK's automatic retries, just burns time, so read the error message and alert a human.

Rate limit headers

For the rate-limit kind, the rate limits guide lists response headers that show where you stand: x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests and the matching -tokens and -project-tokens variants. Follow Retry-After when it's present, and add jitter if you retry yourself. For scale, the gpt-6.1-sol model page lists Tier 1 limits of 500 requests per minute and 500,000 tokens per minute.

Tool call options

tool_choice takes auto, required, a specific function, or none, and parallel_tool_calls: false limits the model to one call per turn. The function calling guide also shows a second pattern for the loop in step 7: append the previous output items to your own input list instead of chaining with previous_response_id.

Conversation state options

You have three options, and they differ in billing and in how long your data lives.

OptionHowWhat to know
previous_response_idPass previous_response_id=response.id on the next callDoesn't carry over top-level instructions, so resend them. Every prior input token in the chain is billed as input again
Conversations APIclient.conversations.create(), then pass conversation=conversation.idItems persist indefinitely, and /v1/conversations data is kept until deleted
Manual replayYou keep the items and send them in input yourselfYou decide what's kept, at the cost of writing it

The details come from the conversation state guide and OpenAI's Chat Completions migration guide. That guide also shows the two moves for anyone coming from Chat Completions: messages becomes input plus instructions, and response_format becomes text.format. Tutorials still teaching older model families are stale on the model ID, so check it before you copy their code.

For illustration only, not part of this project, here's the Conversations API route. It's the better fit when a conversation outlives a request chain:

Python
conversation = client.conversations.create()

response = client.responses.create(
    model="gpt-6.1-sol",
    input="The invoice page crashes when I click download.",
    conversation=conversation.id,
)

Retention matters as much as billing. Per the data controls guide, responses are stored by default for 30 days, and store: false turns that off. Conversations data stays until you delete it and isn't eligible for zero data retention (ZDR). With ZDR on, store is forced to false. If tickets contain personal data, decide which option you're on before launch, not after.

Choosing a model and what it costs

Four levers keep the bill bounded: the model, max_output_tokens, the cached-input price, and the hard limit from step 2. For a pure high-volume classifier, OpenAI's models page points to gpt-6-luna; this article doesn't price it, so check that page. Prices below are per million tokens, from the model pages, checked 2026-10-01:

ModelInputCached inputOutput
gpt-6.1-sol$2$0.10$10
gpt-6-astra$10$1$50

Both have a 1,050,000-token context window and 128,000 maximum output. Both rows come from OpenAI's model pages, checked 2026-10-01: OpenAI's Astra model page and OpenAI's Sol model page. Prompts over 272K input tokens cost 2x on input and cached rates and 1.5x on output for the whole request. The arithmetic below assumes a short ticket.

Here's one triage request as plain arithmetic, not a measured figure. With 1,000 input tokens and 300 output tokens on Sol, that's 1,000 x $2 / 1M plus 300 x $10 / 1M, or $0.002 + $0.003 = $0.005. The same request on Astra is $0.01 + $0.015 = $0.025, five times as much. The docs call Sol near-Astra performance at a lower cost, which is why a classifier starts there. If Sol misclassifies some tickets, try Astra on those and use this table to see whether the gap is worth five times the cost. Our GPT-6.1 Sol pricing comparison goes deeper on the trade.

Treat max_output_tokens as a ceiling, not a target, and keep it above what a normal answer needs (see step 5). Read the usage object on a finished response to see what each request actually consumed, and compare it with your arithmetic before you scale up. Chained state from step 8 adds to the input side, so a long previous_response_id chain is the most common reason a real bill beats the estimate. The hard limit from step 2 is the backstop. Raise it when you have the traffic data to justify it.

Ship checklist

Before this goes to real users, check every line. Each traces back to a step above.

  • The key lives only in server environment variables or a secrets manager, never in frontend code or source control.
  • .env is in .gitignore, and each environment has its own key.
  • Your endpoint authenticates the caller, caps input length and rate limits per user.
  • A hard spend limit is set, not only an alert.
  • Every request sets max_output_tokens, with headroom above a normal answer.
  • You check response.status for incomplete and the output items for a refusal before trusting a result.
  • The tool loop has a round cap, and lookup_order validates what the model passes it.
  • instructions are resent on every chained call.
  • You've decided on store, retention and ZDR, since responses are kept 30 days by default.
  • You treat spend-limit and credit 429s as alerts, not retries, and log _request_id on every failure.
  • You pin gpt-6.1-sol and openai>=3.22,<4, and re-run your tests when either changes.

Common mistakes

Most failures in this tutorial come from a short list.

  • Adding a route below app.mount(...). It answers 405, because the mount at / matches first. Paste every route above the mount.
  • Reading output_text and nothing else. A response can be incomplete or contain a refusal, so check status and the output items.
  • Putting the key in frontend code. Keep it on the server and call your own endpoint.
  • Retrying a spend-limit or credit 429. Waiting a few seconds doesn't fix it, so alert a person.
  • Forgetting that previous_response_id drops the top-level instructions, and that the chain re-bills every prior input token. Resend the instructions on every chained call.
  • Running a tool loop with no round cap, or trusting the arguments the model passes to your function.

Next steps

  • Work through the Astra migration checklist from step 7 if you're moving tool-calling code off Chat Completions.
  • Swap the in-memory order dictionary for your real order service, and validate order_id before you query anything.
  • Try what this build left out: the Conversations API from the Reference section, zodTextFormat from openai/helpers/zod if you're on Node, and built-in tools such as web search, which the docs cover. The Agents API and multi-agent features are out of scope here.
  • Compare this build with how to use the Claude API, which covers the same shape of feature on a different vendor, if you're choosing between the two.

FAQ

Is the Responses API replacing Chat Completions?

The docs we checked say that for gpt-6-astra and gpt-6.1-sol, you should use the Responses API for tool calling, and that Chat Completions supports requests without tools. That's a hard requirement for tool use on those models. For anything beyond them, check OpenAI's migration guide rather than assuming.

Do I need the Responses API for tool calling on GPT-6?

Yes, on the two models above. The latest-model guide says to use the Responses API for tool calling. Step 7 shows the loop, including function_call_output and call_id.

What's the difference between previous_response_id and the Conversations API?

previous_response_id chains one response to the last, doesn't carry over top-level instructions, and re-bills every prior input token in the chain. The Conversations API stores items in a conversation object that persists indefinitely, and its data is kept until you delete it. The Reference section has the comparison table.

How are chained tokens billed?

All prior input tokens in a previous_response_id chain are billed as input on each new call. Model pricing is per million tokens. On gpt-6.1-sol that's $2 input, $0.10 cached input and $10 output, so long chains cost more than single requests.

Can I call the Responses API from the browser?

Not safely. The Node SDK's dangerouslyAllowBrowser option exists, and its README warns that it exposes your secret API credentials. Route browser calls through your own backend, as in step 4.

Share this article

Author Image

HighCircl Editorial Team

The HighCircl editorial team writes about hiring software engineers, nearshore development, and engineering team building. Our articles draw on direct experience sourcing and placing senior developers across Poland, Hungary, Slovakia, Serbia, Slovenia, Romania, and Spain — and on candid conversations with the CTOs and engineering leads who hire them.

HighCircl is a nearshore engineering network that delivers matched candidate shortlists in 72 hours. Every piece of content we publish is informed by real engagement data: actual developer rates, real hiring timelines, and what separates engineering teams that scale cleanly from those that stall.

Take Me to the Experts

Access our network of industry-leading software engineers.

Start Now