Kimi K3 API pricing: open weights at exactly Sonnet 5's rate
By LW Forge โ maintainer of LLM Scout ยท Updated August 17, 2026
Open-weights models are supposed to be the cheap option. Kimi K3 isn't. Moonshot AI released it on July 16, 2026 as the first open model at 2.8 trillion parameters, and priced the hosted API at $3.00 per million input tokens and $15.00 per million output โ the exact list rate of Claude Sonnet 5, to the cent. That coincidence is the whole story: K3 is not competing on price with the open-weights field, it is competing on capability with the closed flagships, and asking to be paid accordingly.
The rates, as of August 4, 2026
| Model | Input / 1M | Output / 1M | Cache read | Context |
|---|---|---|---|---|
| Kimi K3 (Moonshot direct) | $3.00 | $15.00 | $0.30 | 1M |
| Kimi K3 (OpenRouter) | ~$2.90 | ~$14.00 | varies by host | 1M |
| Claude Opus 4.8 | $5.00 | $25.00 | $0.50 | 200k |
| Claude Sonnet 5 | $3.00 | $15.00 | $0.30 | 200k |
| GLM-5.2 | $1.40 | $4.40 | $0.26 | 1M |
Note that even the cache rate matches Sonnet's โ $0.30, a flat 90% off input, the same discount structure Anthropic uses. And unlike Gemini, K3 charges the same rate across the full million-token window; there is no long-context surcharge waiting at the 200k mark.
What the price difference actually buys
Same workload for both posts in this pair: 300M input tokens and 40M output tokens a month, roughly a mid-sized agentic assistant.
| Model | Input cost | Output cost | Monthly total |
|---|---|---|---|
| GLM-5.2 | $420 | $176 | $596 |
| Claude Sonnet 5 (promo) | $600 | $400 | $1,000 |
| Kimi K3 | $900 | $600 | $1,500 |
| Claude Sonnet 5 (list) | $900 | $600 | $1,500 |
| Claude Opus 4.8 | $1,500 | $1,000 | $2,500 |
Three things fall out of that table. Against Opus 4.8, K3 saves $1,000 a month โ 40%, which is a real number if K3 clears your quality bar. Against Sonnet 5 at list, the saving is exactly zero; there is no cost argument, only a capability and portability argument. And through August 31, 2026, Sonnet 5's launch promotion of $2/$10 makes Claude's workhorse cheaper than the open model โ $1,000 against K3's $1,500. Price your own volumes in the Claude cost calculator, and remember the promo expires while the workload doesn't.
Is the 1M window worth paying for?
Claude tops out at 200k tokens; K3 takes 1M. The instinct is that this saves money on long documents. It mostly doesn't, and it's worth seeing why.
Say you need to reason over a 900k-token corpus. On K3 that's one call: 900k ร $3/M = $2.70 of input. On Claude you chunk it โ five passes of 180k, each carrying the same 30k instruction and schema prefix. That's 900k of document plus 150k of repeated prefix, then a reduce pass over the five summaries. Call it 1.06M input tokens, about $3.20 on Sonnet. The single-call version is cheaper, but by 15%, not by an order of magnitude โ and prompt caching on the repeated prefix erases most of even that gap.
The real argument for 1M isn't the invoice, it's what the chunked version can't do: notice that a clause on page 40 contradicts a table on page 380. Map-reduce loses cross-references by construction, and no amount of prompt engineering puts them back. If your task genuinely needs whole-corpus reasoning, the window is worth paying for; if it's "summarise each section," it isn't. Check what your documents actually measure in the context window calculator before assuming you need the extra room โ most workloads that feel enormous land under 200k.
Where K3 earns the premium, and where it doesn't
The honest case for K3 over Sonnet 5 at an identical price is portability. The weights are public under a modified MIT licence, so the hosted API is a convenience rather than a dependency. You can move to a different host, run your own inference for a workload that can't leave your infrastructure, or keep serving a frozen version after the vendor deprecates it. None of that is available on Claude at any price. If a data-residency clause or a vendor-risk review is what's blocking your project, that is worth $0 of price difference and a lot of unblocked work.
The case against is the usual open-weights tax: SDK maturity, structured-output edge cases, observability integrations, and the fact that on aggregators the host serving your request may run a different quantisation than the one you benchmarked. Moonshot's hosted API also runs under Chinese jurisdiction, which is a governance question the open weights answer only if you actually self-host.
Against Opus 4.8 the calculus is different: you're paying 40% less for a model that trades Opus's peak reasoning for a 5ร larger window and open weights. Whether that's a good trade depends entirely on how often your traffic hits the hardest 10% of tasks, where flagship reasoning still wins measurable head-to-heads. Opus vs Sonnet vs Haiku covers how to classify that traffic; the same routing logic in Sonnet 5 vs Haiku 4.5 applies here โ run the cheap tier first, escalate what fails a self-check.
How to decide in an afternoon
Pull 200 real requests. Run them through K3, Sonnet 5 and Opus 4.8 with identical instructions and output caps. Score correctness the way you already do, then compute cost per successful task, charging retries to whichever model caused them. If K3 matches Sonnet, the deciding factor is portability, not price โ and you should make that call deliberately rather than by default. If K3 matches Opus, you've found your $1,000 a month.
Two levers to apply before you conclude anything: cache the stable prefix at 10% of input on all three models (prompt caching), and check whether K3's always-on reasoning is inflating your output token counts, since output bills at 5ร input on every model in this table. If $3/$15 turns out to be more than your task needs, GLM-5.2 at $1.40/$4.40 is the same open-weights bet at less than half the price.