Joan Comadran
Field notes

Engineering

Claude vs GPT for AI agents: what I actually ship with (2026)

Claude vs GPT for building AI agents, as of October 2026: the price tiers now mirror each other almost to the cent, so the choice comes down to API shape, caching and hosted tools. My verdict, a comparison table, and the production code behind it.

October 7, 202612 min read

Here's the verdict, so you can leave early if that's all you came for: I build my AI agents on Claude, and I'd still tell most people that the honest answer to "Claude vs GPT for agents" in late 2026 is "it matters less than you think." As of October 2026 the two price ladders mirror each other almost to the cent, both have 1M-token context windows, both enforce tool schemas strictly, and both host web search for you. What's left to choose on is API shape, how much control you get over caching, which hosted tools you need, and which one your evals prefer on your tool schemas.

If you want a default: pick Claude if your agent lives on long, growing context and you care about controlling exactly what gets cached. Pick GPT if you're building a voice agent, want hosted file search over your documents, or your company already runs on OpenAI or Azure contracts.

One disclosure before the table, because comparison posts rarely make it: every agent I run in production is on Claude. I've used OpenAI's API, but I don't have a shipped GPT agent to compare like-for-like, so the GPT column below comes from OpenAI's docs and pricing pages, not from war stories. The Claude column comes from code I can show you.

Claude vs GPT at a glance (as of October 2026)

All prices are list prices per million tokens (input / output), taken from the Anthropic and OpenAI pricing pages in October 2026. These change often; check before you budget.

Claude (Anthropic)GPT (OpenAI)
Current lineupFable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna
Top tierFable 5.1: $10 / $50GPT-6 Astra: $10 / $50
Between top and midOpus 5.5: $4 / $20—
Workhorse tierSonnet 5.5: $2 / $10GPT-6.1 Sol: $2 / $10
Cheap, high-volume tierHaiku 5.5: from $0.10 / $0.50 (prompts up to 100K tokens)GPT-6 Luna: from $0.10 / $0.50 (short context)
Context window1M tokens on all current models1.05M tokens on the flagship three
Long-context billingFull 1M at the standard rate (except Haiku 5.5 above 100K)Separate, higher long-context rate
Max output128K tokens128K tokens
API you build agents onMessages API (tool_use / tool_result blocks)Responses API (required for tool calling on Astra and Sol)
Schema-enforced tool callsstrict: true on tool definitionsstrict: true; Responses tries to make tools strict by default
Prompt cachingExplicit cache_control breakpoints or automatic; cache reads at 10% of input or lessDiscounted cached input (5–10% of input on the flagship three)
Batch discount50%50%
Hosted toolsWeb search ($10 / 1K searches), web fetch, code execution, tool search, MCP connector, computer and browser useWeb search ($10 / 1K calls), file search, hosted shell / code interpreter, computer use, MCP, tool search
Voice / realtimeText and image in, text outDedicated realtime and voice models

Look at the price rows for a second. Not long ago, a "Claude vs GPT" post was half about cost. Now the cheap tiers start at the same price, the workhorse tiers are identical, and the top tiers are identical. Rate cards aren't where the difference is anymore.

Where they actually differ for agent builders

The API shape

On Claude, an agent turn is a Messages API call. If the model wants a tool, the response has stop_reason: "tool_use" and one or more tool_use blocks; you run them and send back a user turn full of tool_result blocks. On GPT, the newest models require the Responses API for tool calling: the output is a list of items, tool calls come back as function_call items, and you return function_call_output items, optionally chaining with previous_response_id.

Neither is better in the abstract. They are different enough that "we'll just swap the model later" is a week of work, not a config change, once you've built retries, caching, cost logging and error handling around one of them.

Caching control

This is the one that decided it for me. Claude lets you mark exactly where the cacheable prefix ends with cache_control breakpoints (or let the API place them automatically). In a conversational agent, where the system prompt is static and the history grows every turn, being able to say "cache everything up to here" is the difference between cost that grows with the square of the conversation and cost that grows roughly linearly. I wrote up the full story in keeping Claude's token bill from ruining your month.

OpenAI discounts cached input heavily too, so this isn't "Claude caches, GPT doesn't." It's about how much you get to steer it.

Long context billing

Both vendors offer roughly a million tokens of context. The difference is the bill: Anthropic charges the full 1M window at the standard rate on its current models (Haiku 5.5 aside), while OpenAI lists a separate, higher long-context price on Astra, Sol and Luna. If your agent stuffs whole codebases, contracts or long transcripts into context, model that difference before you pick. If your prompts are a few thousand tokens, it doesn't matter.

Hosted tools

Both now run web search on their side and support MCP. OpenAI has hosted file search with managed storage and a full line of realtime voice models; if you're building a phone or voice agent, that alone probably decides it. Anthropic has server-side web search and web fetch, code execution, a tool-search tool for large tool sets, and an MCP connector that talks to remote MCP servers straight from the Messages API.

What I actually run in production (and why)

None of my agents use LangChain or the Vercel AI SDK. Every one calls @anthropic-ai/sdk directly, because the provider-specific features are exactly the ones I'm relying on. Here are two, straight from the repos.

Claient: a multi-tenant WhatsApp agent on Haiku

Claient answers and qualifies WhatsApp leads for small businesses (the build story is here). The model choice is gated by who pays:

ts
// With the platform key (we pay), only Haiku is served.
// Tenants who bring their own key (BYOK) can pick Sonnet or Opus.
export const MODELO_BASE = "claude-haiku-4-5-20251001";

export function modeloEfectivo(opts: { apiKey?: string | null; model?: string | null }) {
  if (!opts.apiKey) return MODELO_BASE;
  return opts.model || getClaudeConfig().model;
}

The reply runs in a tool-use loop with three booking tools (check availability, book, reschedule). Two guards matter more than the model: a hard cap of five loop iterations, and a time budget, because a serverless function killed mid-loop means the lead got a message but the webhook didn't finish, and Meta's retry sends it twice:

ts
const MAX_ITER_TOOLS = 5;
const MARGEN_VUELTA_MS = 20_000; // SDK timeout 15 s + 1 retry must fit

for (let i = 0; i < MAX_ITER_TOOLS; i++) {
  if (i > 0 && !hayMargen(opts.deadline)) break; // fall back instead of dying mid-turn
  const respuesta = await claude.messages.create({ model, max_tokens, system, messages, tools });
  if (respuesta.stop_reason === "tool_use") {
    // run each tool_use block, push the tool_result blocks, loop again
    continue;
  }
  return textOf(respuesta);
}

Lead qualification is a separate call that forces a tool, tool_choice: { type: "tool", name: "registrar_lead" }, with nullable fields and an enum for lead temperature (cold, warm, hot). Forcing a tool is the most reliable way I know to get structured data out of a conversation, and GPT can do the same thing with tool_choice plus strict mode. The cache breakpoints sit on the system prompt and the latest turn, and the history sent is a window of the last 16 messages plus a rolling summary. Every call logs input, output, cache-read and cache-write tokens and their USD cost per tenant, with a monthly spend cap.

An honest footnote: Claient still pins Haiku 4.5, which is now a legacy model. Haiku 5.5 lists at a tenth of the price. That's the obvious thing to test next, and it's also a good example of why you want evals: "cheaper and newer" is a hypothesis until your own conversations say otherwise.

Pollitics: two models, two jobs

For pollitics.es, a weekly job drafts candidate questions about the week's news for a human to review. It's split into two Claude calls on purpose:

  1. Research on Sonnet 5.5 with Anthropic's server-side web search (max_uses: 10, effort medium). Most of the cost here is input tokens from the pages it reads, so the cheaper model does the reading.
  2. Drafting on Opus 5.5 with JSON-schema structured output (effort high), where editorial nuance matters and the input is small.

It's two calls, not one, because structured output doesn't carry the search citations, so the research comes back as text with URLs and the drafter works from that. The code also handles pause_turn (the server-side search loop paused, send it back to continue) and refusal stop reasons, and the whole thing has to fit inside a 150-second Edge Function. A deterministic review step then flags anything that breaks the editorial rules before a person approves it.

So why Claude?

Not because GPT couldn't run either of these. The reasons are concrete and a little boring:

  • Cache breakpoints matched the architecture. Static script plus growing history is exactly the shape explicit cache_control is good at.
  • Tiering within one vendor. Haiku for every message, Sonnet for reading the web, Opus for the writing that matters. One SDK, one usage format, one cost table. Dairector's script doctor runs on Claude too.
  • BYOK stays simple. Tenants paste one kind of key. Supporting two providers' keys, models and error classes would double the surface for very little customer value.
  • Compounding. Timeouts, retries, error classification (AuthenticationError vs transient 429/5xx), cost logging: all written once, against one SDK. That's the real lock-in, and it's mine, not the vendor's.

When to use which

Choose Claude when:

  • Your agent is conversational or long-running and context grows every turn, so caching control pays for itself.
  • You need very long context and don't want a long-context surcharge on the mid and top tiers.
  • You want one vendor with a clear cheap / workhorse / heavy ladder to route tasks across.

Choose GPT when:

  • You're building a voice or phone agent and want first-party realtime models.
  • You want hosted file search over your documents without running your own vector store.
  • Your organisation already has OpenAI or Azure OpenAI contracts, data agreements and monitoring in place. That's a perfectly good reason; procurement is engineering too.

Either way:

  • Build a small eval set from real conversations (twenty is enough to start) and run both on your actual tool schemas before you commit.
  • Put the model call behind one module of your own so a future switch touches one file.
  • Cap loop iterations and spend from day one. Neither vendor will do that for you.

If you're weighing this for a real product and want someone who has shipped agents end-to-end to make the call with you, that's what I do. For the rest of the stack around the model, see how I ship an AI product in a weekend.

FAQ

Is Claude or GPT better for building AI agents in 2026?

Neither wins across the board. As of October 2026 the price tiers mirror each other (Sonnet 5.5 and GPT-6.1 Sol are both $2/$10 per million tokens; Haiku 5.5 and GPT-6 Luna both start at $0.10/$0.50), both offer 1M-token context, 128K max output, strict schema-enforced tool calls and hosted web search. Claude is the better default if you want explicit control over prompt caching and long context billed at the standard rate; GPT is the better default for voice agents, hosted file search, or teams already on OpenAI or Azure. Run your own eval on your own tool schemas before committing.

Is Claude more expensive than GPT for agents?

Not anymore at list price. As of October 2026 the mid tiers (Claude Sonnet 5.5 and GPT-6.1 Sol) are both $2 input / $10 output per million tokens, the cheap tiers (Claude Haiku 5.5 and GPT-6 Luna) both start at $0.10 / $0.50, and the top tiers (Claude Fable 5.1 and GPT-6 Astra) are both $10 / $50. Claude Opus 5.5 sits in between at $4 / $20. Both offer 50% off for batch work and discounted cached input. Real cost differences come from how much context you resend and how well you cache it, not the rate card.

Can I switch an agent from Claude to GPT later?

Yes, but it is not a one-line change. The Claude Messages API uses tool_use and tool_result content blocks; OpenAI's Responses API, which GPT-6 Astra and GPT-6.1 Sol require for tool calling, uses function_call and function_call_output items. Caching, usage fields, error classes and hosted tools also differ. Keep the model call behind one small module of your own so a switch touches one file, and keep an eval set so you can tell whether the switch was an improvement.

Which model does Joan Comadran use for production AI agents?

Claude, called directly through the Anthropic TypeScript SDK with no agent framework. Claient, the WhatsApp lead agent, runs Claude Haiku 4.5 by default with tool use for appointment booking and forced tool calls for lead extraction, and lets tenants bring their own key to use Sonnet or Opus. The weekly-question pipeline for pollitics.es uses Claude Sonnet 5.5 with server-side web search for research and Claude Opus 5.5 with JSON-schema structured output for drafting.

Do I need LangChain or another framework to build an agent with Claude or GPT?

No. A tool-use agent is a loop: call the model with tool definitions, execute any tool calls it returns, send the results back, and stop when it answers in text. Both official SDKs make that a few dozen lines, and both vendors now offer optional helpers (Anthropic's Tool Runner, OpenAI's Agents SDK). A framework can help with many providers or complex orchestration, but it also hides the provider-specific features, like cache breakpoints and server tools, that most of the cost and reliability wins come from.

Keep reading

03