9 min read B2C power user

Claude Sonnet 5.5 Pricing: Same Rate, 30% Less Per Task

Claude Sonnet 5.5 keeps Sonnet 5 list pricing but needs far fewer tokens per task. Here is what the new cost curve means for your monthly AI bill.

Claude Sonnet 5.5 Pricing: Same Rate, 30% Less Per Task

Search interest in claude sonnet pricing sits around 720 queries a month with a cost per click near $7, and it climbed roughly 32 percent over the last three months. On September 28, 2026, Anthropic gave that audience a genuinely unusual answer. Claude Sonnet 5.5 ships at exactly the same list price as Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. The headline is not a discount. The headline is that the same task now consumes far fewer tokens, and Anthropic measures the difference at up to 30 percent less cost per task.

That distinction matters more than any price cut. A lower per-token rate is easy to model. A model that reaches the same answer in fewer steps changes the shape of your bill, because it attacks the part of agentic spend that actually scales: the number of round trips, not the price of each one.

What actually changed between Sonnet 5 and Sonnet 5.5

The benchmark movement is the tell. On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scores 70.6 percent against Sonnet 5’s 10.3 percent. That is not a tuning pass. It is a different class of model wearing the same price tag. Sonnet 5.5 also generates output more than 30 percent faster, which matters to anyone running long agent sessions where wall-clock time is the real constraint.

The pricing page frames this as a cost-per-task curve rather than a rate card, and the curve is where the value lives. At Medium effort, which is the default in the Claude apps, Sonnet 5.5 beats Sonnet 5’s best score for under a tenth of the cost per task. On CursorBench 4.0, Sonnet 5.5 at Low effort exceeds Sonnet 5’s best score for less than a tenth of the cost per task. On AA-Briefcase, a long-horizon knowledge work benchmark, Medium effort bests Sonnet 5’s ceiling at roughly one ninth of the cost.

The practical reading for a heavy user is simple: if you were using Sonnet 5 at High or Max effort to squeeze quality out of it, you can now drop to Low or Medium and get a better result for a fraction of the spend. Effort level is a cost dial, and the new model moved the whole curve to the left.

Claude model pricing: where Sonnet 5.5 sits in the lineup

Sonnet 5.5 is the second model in the Claude 5.5 family, positioned deliberately below Claude Opus 5.5. Anthropic is explicit about the split. Opus 5.5 is for complex, open-ended work that needs sustained judgment. Sonnet 5.5 is strongest at well-scoped everyday tasks: fixing bugs, and creating polished documents, slides, and spreadsheets.

The numbers back the tiering. On GDPval-AA, a test of real-world work across occupations, Sonnet 5.5 scores 1844 against Opus 5.5’s 1846. That is effectively a tie on knowledge work. But on FrontierCode 1.1, Opus 5.5 leads 54.4 to 46.2 percent, and on Humanity’s Last Exam Opus 5.5 holds 67.7 percent with tools against Sonnet 5.5’s 64.5 percent. Sonnet 5.5 is close on the work you do every day and clearly behind on the work that requires deep reasoning chains.

That is a routing decision, not a quality debate. The expensive mistake heavy users make is defaulting to the top-tier model for tasks that a mid-tier model now handles at a tenth of the cost per task. Sonnet 5.5 is priced the same as Sonnet 5, which means the correct move is to demote Sonnet 5 entirely and move anything you were running on Opus 5.5 that is genuinely well-scoped down to Sonnet 5.5.

Anthropic also confirmed that Claude Haiku 5.5, built for high-volume and cost-sensitive applications, joins the Claude 5.5 family in the coming weeks. If you run batch classification, extraction, or cheap sub-agents, that is the tier to wait for rather than over-spending on Sonnet today.

Claude token pricing and the cross-provider comparison

TokenKarma does not cover Claude in isolation, and this release is worth measuring against the rest of the field. On FrontierCode, Anthropic reports that Sonnet 5.5 at High effort matches GPT-6 Sol’s best score for about a fifth of the cost per task. On CursorBench, Sonnet 5.5’s best score lands within roughly two points of Opus 5.5 while sitting near GPT-6 Sol’s 52.1 percent.

The external tester quotes point the same direction. CodeRabbit reported that Sonnet 5.5 shows better judgment than Sonnet 5 while spending significantly fewer output tokens, and specifically that Sonnet 5’s habit of reaching for web search too often is gone. That is a cost story disguised as a quality note: runaway tool calls are one of the biggest sources of wasted spend in agent workflows. Base44 ran 118 real app builds and found Sonnet 5.5 matched Opus 5 quality in 3.6 iterations on average, where Opus 5 needed 7.7, with the fewest failed tool calls of any model they compared.

Fewer iterations and fewer failed tool calls compound. Each avoided retry is a saved context re-send, and each saved re-send is output tokens you do not pay for. This is the same lesson the Nvidia SoL-Pi harness research made last week: the biggest lever on agent cost is doing less work per task, not paying less per token.

Claude code token cost: what to change today

If you pay for Claude Code, the API, or a Max plan, the release is actionable right now rather than after a migration project.

Demote your default model. Anything you run on Sonnet 5 should move to Sonnet 5.5. It is the same price with dramatically higher agentic accuracy, which means fewer retries on the tasks that were quietly failing.

Lower your effort setting before you lower your quality expectations. The cost-per-task charts show Sonnet 5.5 at Low or Medium effort beating Sonnet 5’s best score. Run a week at Medium where you previously used High and compare the cost per solved task, not the cost per token.

Re-audit what you route to Opus 5.5. Well-scoped bug fixes, document generation, slide and spreadsheet work, and code review are now Sonnet 5.5 territory. Keep Opus 5.5 for genuinely open-ended, long-horizon judgment work.

Watch your cache read ratio. Cache reads stay at $0.20 per million tokens, a 50x discount against output. A model that batches tool calls and trims redundant steps will reuse cached context more efficiently. If your cache read share drops after switching, your workflow is re-sending context it should be caching.

Measure cost per solved task, not cost per token. The Sonnet 5.5 launch is the clearest proof yet that per-token rates are a poor proxy for what you actually spend. Instrument the number of attempts a task takes and multiply.

What the Sonnet 5.5 launch means for your budget

Anthropic shipped a model that costs the same on paper and up to 30 percent less in practice, with a benchmark jump so large it effectively resets the definition of a mid-tier model. The companies that benefit most are the ones that treat model choice as a routing problem with a measurable cost per solved task, and the ones that lose are the ones that leave a top-tier default in place out of habit.

Run the switch deliberately. Demote Sonnet 5, drop your effort defaults one notch where the charts support it, move well-scoped work off Opus 5.5, and hold Haiku 5.5 in reserve for the high-volume tier when it lands. Then measure cost per solved task across Claude, GPT-6, and Gemini on your own workloads rather than trusting a rate card. The published price did not move, and your bill should.