โ† All articles

Claude Sonnet 5's Launch Price Is Now Permanent

By LW Forge โ€” maintainer of LLM Scout ยท Updated August 12, 2026

When Claude Sonnet 5 launched, its $2/$10 per million token rate carried an expiration date: standard pricing of $3/$15 was scheduled to take effect on September 1, 2026. Anthropic has since confirmed that increase will not happen โ€” the introductory price is now the permanent price. If you built a budget, a migration plan, or an argument for switching to Haiku around that September deadline, the math underneath it just changed.

What actually changed

Anthropic's official pricing page now states plainly that "the $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur." This isn't a rumor or a third-party estimate โ€” it's the model pricing table itself, the same source this site's Claude cost calculator treats as ground truth.

Nothing else about Sonnet 5 moved. The 200k context window, the cache-write multiplier (1.25x for a 5-minute cache, 2x for a 1-hour cache), and the 10%-of-input cache-read rate are unchanged โ€” only the base input and output rate stopped being temporary.

The dollar impact, worked

Take a mid-size product sending 50M input tokens and 10M output tokens a month through Sonnet 5 โ€” a reasonable size for a chat feature with a few thousand active users.

ScenarioInput costOutput costMonthly total
Rate that was scheduled ($3 / $15)$150$150$300
Rate that actually applies ($2 / $10)$100$100$200

That's a permanent $100-a-month gap at this volume, and it scales linearly: a workload ten times this size keeps $1,000 a month that a September 1 budget line would have written off. If you'd already padded a forecast for the increase, the padding is now just margin. If you hadn't โ€” because the promo period felt easy to forget until the invoice changed โ€” this removes a bill shock that was two weeks away.

It also resets the Haiku comparison

Sonnet and Haiku 4.5 are priced with the same input-to-output ratio, so the gap between them is a clean multiple of whichever Sonnet rate is current. At the scheduled $3/$15, Sonnet cost exactly 3x Haiku. At the now-permanent $2/$10, it costs exactly 2x. That's a real, structural narrowing, not a temporary one: the case for defaulting to Haiku purely to dodge Sonnet's price got noticeably weaker, because the tier you're avoiding is now cheaper than it was ever going to standardly be. The worked pipeline math is in Sonnet 5 vs Haiku 4.5: when the cheaper Claude is enough, updated to the current rate.

The same shift touches the three-tier picture. Opus 4.8, Sonnet 5 and Haiku 4.5 no longer scale by a tidy 5:3:1 โ€” it's 5:2:1 now, which changes how much a tier upgrade actually costs you at the margin. See Opus vs Sonnet vs Haiku for the full table.

The context window detail most people missed

Buried in the same pricing update is a second change worth knowing about: Claude 4.6 and later models, Sonnet 5 included, now bill the full 1M-token context window at standard per-token rates โ€” there's no long-context surcharge for requests over 200k tokens, the way some competing APIs apply a higher rate past a threshold. A 900,000-token request costs the same per token as a 9,000-token one. If your use case involves large documents or long agent transcripts, run the real size through the context window calculator to see what that means for your specific request shape, and context window in practice for how much actually fits before quality starts to degrade.

What to actually do about it

If you were planning to switch workloads to Haiku on September 1 specifically to avoid the price bump, re-run the comparison โ€” the 2x gap at permanent pricing may no longer clear the bar that made switching worth the engineering effort. If you'd forecast Sonnet spend at $3/$15 for anything beyond this month, correct the model now rather than discovering the overestimate in a quarterly review. And if you haven't priced your actual traffic yet, don't extrapolate from either number in this article โ€” put your real token counts into the Claude cost calculator, which pulls live pricing rather than a static table, and compare against GPT-5.6 if you're still deciding between providers rather than tiers.

Anthropic didn't have to make this permanent โ€” an introductory rate reverting on schedule would have been unremarkable. Making it standard, weeks after the fact and with the increase already telegraphed, reads like a response to how competitive the budget-to-mid tier of the market has become. Whatever the reason, the number that matters is the one now printed on the pricing page, not the one that was scheduled.