Claude API pricing looks simple on the surface — a per-million-token rate for input, another for output — but the effective price you pay depends on which model you pick, whether you use prompt caching, whether you can wait for a batch, and how much of your traffic can move to a cheaper model. This piece walks through all of it with real dollar figures, then answers the two questions most people arrive with: how does the API compare to a Claude Pro subscription, and how do I stop overpaying.
The price list
Every Claude model is priced per million tokens, separately for input and output. The current lineup:
- Claude Haiku 4.5 — $1 input, $5 output per 1M tokens. 200K context window. Small, fast, cheap.
- Claude Sonnet 5 — $2 input, $10 output per 1M tokens. 1M context window. The current default for most workloads.
- Claude Sonnet 4.6 — $3 input, $15 output per 1M tokens. 1M context. Older Sonnet generation.
- Claude Opus 5 (and 4.6 / 4.7 / 4.8) — $5 input, $25 output per 1M tokens. 1M context. For hard reasoning and long-horizon agentic work.
- Claude Fable 5 and Fable 5.1 — $10 input, $50 output per 1M tokens. 1M context. Anthropic's most capable widely-released model.
Output tokens cost five times as much as input tokens on every model. That's the single most important thing to internalize about the pricing sheet: what you send in is cheap, what Claude writes back is expensive. Almost every optimization below comes back to that ratio.
What actually counts as an input token
Every byte you send counts: the user's message, the entire conversation history you resend on each turn, the system prompt, tool definitions, tool results, and any files or images you attach. If you're running a chat app and you send the last twelve turns back every request, you pay input-token cost on all of them, every time. Multiply by number of requests and it adds up fast — but that's exactly where caching becomes the biggest lever on your bill.
Prompt caching: the 90% discount most teams don't use
Prompt caching is Anthropic's most under-used cost tool. Mark the stable prefix of a request as cacheable, and subsequent requests that share that exact prefix pay roughly one-tenth of the normal input-token rate for it. On a large system prompt, a rag context, or a repeated tool schema, this is often 80–95% of your input token spend gone.
The rules are strict but simple. Caching matches a prefix — any single byte that changes anywhere in that prefix invalidates the whole cache. The render order is `tools`, then `system`, then `messages`. So the pattern is: put stable content first (frozen system prompt, deterministic tool list), volatile content last (the user's actual question, timestamps, per-request IDs).
First request pays a slight write premium (about 1.25x normal input cost for cached tokens). Every subsequent request that hits the cache pays ~10% of the normal input rate. On a workflow with a 20K-token system prompt and 100 daily requests, that's the difference between paying for 2M input tokens a day and paying for 200K. On Opus 5, that's $10 a day versus $1.
The Batch API: 50% off, if you can wait
If your workflow doesn't need a real-time response — nightly summarization, bulk enrichment, weekly digest generation, backfill jobs — send it through the Batch API. You submit a batch of requests, Anthropic processes them within 24 hours, and every model runs at 50% of its listed rate. Both input and output.
That's a straight halving of every cost number in this article. For any workflow where latency doesn't matter, batching should be the default. The gotcha is that batches complete asynchronously — you poll for status and pull the results when they're ready — so it doesn't fit into a request-response chat UI. But for background jobs it's the biggest cost lever after caching.
Extended thinking (adaptive thinking): a token cost, not a separate charge
Extended thinking — where Claude reasons internally before answering — bills as output tokens at the model's normal output rate. It's not a separate line item. On Opus 5 with adaptive thinking at high effort, expect a hard problem to spend a few thousand thinking tokens on top of the actual answer. That means one hard question can cost $0.10–$0.30 in output tokens on Opus 5, versus $0.02 on Sonnet 5, versus under a cent on Haiku 4.5.
The output-cost multiplier is why effort tuning matters. Every rung up the effort ladder (low → medium → high → xhigh → max) sends more output tokens, and cost scales roughly with output token count. High effort is the sweet spot for coding and long-agent work; low effort is often invisible-quality-loss for classification, extraction, and short chat responses. Measure before setting max as your default.
Real workflows: what teams actually pay a month
Abstract dollars per million tokens don't land. Here are four concrete scenarios with rounded numbers.
- Daily newsletter digest — one workflow run per day, 20K tokens in (RSS content), 3K tokens out (a summary), Sonnet 5, no caching: about $0.07 per run, $2/month. Move it to Haiku 4.5 and it's about $0.03/run, less than $1/month.
- Marketing-ops chat assistant — 500 short conversations a month, 10 turns each, 4K tokens per turn average, Sonnet 5 with prompt caching on a 30K system prompt: roughly $30/month. Without caching, same workload: about $180/month.
- AI email triage for one inbox — 1,000 messages a day scored on Haiku 4.5, 500 tokens in per message: about $15/month, running 24×7.
- AI SDR that researches leads — 200 leads a day, 15K input tokens each (page-scrapes + system prompt), 2K output tokens each, Opus 5 for the reasoning: about $180/month without caching, or $30/month with the research prompt cached.
The pattern in every one of these: model choice sets the ceiling, prompt caching sets whether you actually hit it, and output-token discipline (shorter responses, structured outputs, cache the schema) sets the floor.
Claude API vs Claude Pro / Max / Team: which one should you actually pay for?
This is the question most people are really asking when they search 'Claude API pricing'. A Claude subscription — Pro at $20/month, Max at $100 or $200, Team at $30/user — gives you Claude in the browser and desktop app. The API is the developer surface: no chat UI, just requests, priced per token.
- If you want to talk to Claude in a browser or in Claude Code, you want a subscription. The API can't do that for you.
- If you want Claude to run inside your product, or to power a workflow, or to be callable by another tool, you want the API. A subscription won't do that.
- The break-even is not a fixed dollar amount — it depends on how many output tokens you generate a month. A hundred Sonnet 5 messages of 1,000 output tokens each is roughly $1 on the API but requires no subscription. A million-token research task is $10 on Opus 5, versus a $20/month Pro subscription that would refuse the length.
For most teams the answer is both, for different jobs. A Pro or Max subscription for personal use in the browser (the browsing experience is worth the flat fee). API access for anything that runs inside your product, automates a workflow, or fires from a script. The two aren't substitutes; they're two different products.
Where token markup enters — and how to skip it
Every automation platform in the AI workflow category — Zapier, Make, n8n Cloud, Gumloop, most of the newer entrants — resells Claude API access with markup baked in. The markup ranges from 2x to 5x the raw Anthropic rate. On a workflow that fires 500 times a month with a few Opus 5 calls per run, that's the difference between $30 and $150 a month, on top of the platform's own subscription fee.
The alternative is bring-your-own-key. You supply your Anthropic API key, the automation platform executes Claude calls with it, and you pay Anthropic directly with no middleman markup. This is the model Pilotran runs on: every paid plan lets you plug in your own Anthropic (or OpenAI) key on day one. You still pay the platform subscription — but not a variable AI surcharge on top.
On a 500-run-a-month workflow with Opus 5, the difference between markup and BYO-key over a year is enough to cover several months of the platform itself. It's the biggest hidden line item in the category, and the fastest fix.
The short version
- Haiku 4.5 for triage and high-volume simple tasks ($1/$5). Sonnet 5 for most workloads ($2/$10). Opus 5 for hard reasoning ($5/$25). Fable 5/5.1 for the most demanding work ($10/$50).
- Prompt caching is a 90% discount on any stable prefix. Turn it on before doing anything else.
- Batch API is 50% off for any non-realtime workload. Use it for background jobs.
- Output tokens cost 5x input. Every optimization that shortens output pays back fastest.
- The API and Claude Pro / Max / Team subscriptions are different products. If you're building automation, you want the API. For personal browser use, get the subscription.
- If an automation platform is billing you AI credits on top of a subscription, you're paying markup. Move to a BYO-key platform.