The Rate Card
Claude API pricing: the rate card
| Model | Input | Output | Cache read | Cache write | Context | Max output |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 | $2.50 | 1M tokens | 128K tokens |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | $5.00 | 1M tokens | 128K tokens |
| Claude Haiku 5.5 | — | — | — | — | — | — |
All figures are US dollars per million tokens on Anthropic's direct API. Prices verified September 30, 2026.
Cost Calculator
What your Claude API bill will actually cost
Per-token prices are hard to feel. Fill in three numbers and this page will tell you what a request costs and what a month costs.
At 100,000 requests a month, the support-bot preset runs about $900/month on Sonnet 5.5 and about $1,800/month on Opus 5.5. The cheaper model is highlighted.
These are list-price estimates. They exclude caching, batch discounts and any negotiated rate.
Model Comparison
What Claude API pricing buys you
Sonnet 5.5 is the everyday model. Anthropic positions it for well-scoped tasks: fixing bugs, producing documents, slides and spreadsheets. It scores 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5, and generates output 30%+ faster. If your workload is high-volume and well-defined, this is the one you want.
Opus 5.5 is the judgment model. It's for complex, open-ended work where a wrong call is expensive: codebase-wide migrations, audits, long agent runs. Anthropic says it costs 40% less to run than Opus 5 on typical workloads, mostly because it uses fewer tokens per task rather than because the rate card is lower.
Which one to pick is mostly a question of shape, not quality. If your requests are short and well defined, like extracting fields, classifying tickets or drafting replies, Sonnet 5.5 is the sensible default, and the 2× gap in input and output rates is real money at volume. Opus 5.5 earns its rate when a wrong answer costs more than the tokens: codebase-wide migrations, audits, and long agent runs where one bad decision cascades.
There is one place the two models price identically, and it matters more than people expect. A cache read costs $0.20 on both, so a workload built around one large repeated prefix gets the same discount on either model.
Both models share the same envelope: 1M-token context window. Roughly 750,000 English words, about ten novels, or a large codebase. 128K max output per request. Knowledge cutoff: June 2026.
Thinking is on by default. Sonnet 5.5 uses adaptive thinking with configurable effort (low, medium, high, xhigh, max); the default is high on the Claude Platform and medium in the Claude apps. Opus 5.5's thinking is always on and cannot be disabled; its default effort is medium. Both stay available through at least September 2027. Anthropic commits to no earlier than September 28, 2027 for Sonnet 5.5 and September 22, 2027 for Opus 5.5, with a six-month legacy window after that.
Prompt Caching
How prompt caching changes your Claude API bill
Caching is where most teams find their savings, and where the arithmetic surprises people.
Writing to the cache costs more than normal input. On both models the premium is 25%: $2.50 vs $2.00 on Sonnet 5.5, $5.00 vs $4.00 on Opus 5.5.
Reading from the cache costs a fraction of normal input. On Sonnet 5.5 a cache read is $0.20, one tenth of the $2.00 input price. On Opus 5.5 it is $0.20 against a $4.00 input price, one twentieth.
That asymmetry is the whole point: you pay a 25% premium once to write a prefix, and one tenth to one twentieth of the input price every time you reuse it. Any prefix you send twice has already paid for itself.
Anthropic notes that cache reads make up the majority of agentic and coding work costs, which is why the Opus 5.5 cache-read price was cut 60% from Opus 5.
Caching is not free money, though. It pays when a prefix repeats. If every request carries a fresh document, you pay the 25% write premium and never read it back, which costs more than sending that text as ordinary input. The break-even is easy to hold in your head: a prefix has to be reused at least twice before caching beats paying the input price twice, and a cache entry has a lifetime, so a prefix that shows up once an hour may expire between uses. The workload that wins is the boring one, where the same system prompt, the same instructions and the same reference material go out on every request.
Methodology
How we check Claude API prices
Every number on this page points to the page we read it from, and every row carries the date we last checked it.
Prices are read directly from Anthropic's model announcement pages. We open the page, read the pricing table, and record the date. Specifications (context window, max output, knowledge cutoff, end-of-life commitment, thinking defaults) are read from the official AWS Bedrock model cards for each model.
No aggregator data. If a figure isn't on a page we opened ourselves, it doesn't go in the table. That's why the Haiku 5.5 row is empty. Nothing is scraped or auto-imported. A stale row is worse than a blank one.
Prices are per million tokens and change when Anthropic ships or re-prices a model, not on a schedule. We re-check when a model is announced, and run a full audit of this table every month. The "last checked" date above is the real date of the last audit.
Transparency
How this page makes money
Right now: it doesn't. There are no ads and no affiliate links on this page.
If that changes, the rules are simple. Any commercial link will be labelled, and it will only ever point to something that is genuinely cheaper than paying list price for the same tokens. We won't accept payment to move a number on this page.
This page can't chase scale; the sites that do are bigger and update faster. The only thing it can trade on is being right and showing its work.
FAQ