GPT-6 Astra: what the launch changes and what it costs
Updated September 9, 2026 ยท LW Forge
A practical overview of GPT-6 Astra: its main capabilities, API price, long-context surcharge, rollout status and the teams that should test it first.
Practical guides on AI API pricing, tokens and cost planning. Every article links to the calculator that lets you run the numbers for your own case.
Updated September 9, 2026 ยท LW Forge
A practical overview of GPT-6 Astra: its main capabilities, API price, long-context surcharge, rollout status and the teams that should test it first.
Updated September 7, 2026 ยท LW Forge
Claude Fable 5.1 pricing lands at $10 per million input tokens and $50 output, with cache reads cut 75% to $0.25. Compared to Opus 5, Sonnet 5, and Haiku 4.5, worked in real dollars.
Updated August 31, 2026 ยท LW Forge
Google's new Gemini 3.7 Flash cuts input and output pricing roughly in half versus 3.5 Flash, confirmed on Google's own pricing page โ but the rate is introductory and reverts on January 1, 2027. Here's the real math.
Updated August 31, 2026 ยท LW Forge
Compare GPT-5.6 Sol and GPT-5.6 Terra by token, task, and workload. See when Terra is enough and when Sol can justify its premium.
Updated August 29, 2026 ยท LW Forge
Business Insider and The Information report Nvidia is closing in on buying Hugging Face for around $13B, months after the open-source hub turned down a smaller Nvidia investment. Here's what's confirmed and what isn't.
Updated August 25, 2026 ยท LW Forge
Why reasoning and thinking tokens are billed as expensive output tokens, how hidden chains of thought multiply API costs, and strategies to control your budget.
Updated August 12, 2026 ยท LW Forge
Anthropic confirmed the $2/$10 introductory rate for Sonnet 5 won't revert to $3/$15 on September 1 as originally planned. Here's what actually changes for budgets, the Haiku comparison, and every calculator on this site.
Updated August 12, 2026 ยท LW Forge
OpenAI and Anthropic both cut prices in half for the Batch API, no exceptions and no fine print on the rate. The catch isn't the discount โ it's whether your workload can tolerate a 24-hour wait, and most people size that wrong.
Updated August 12, 2026 ยท LW Forge
Portuguese costs more tokens than English for the exact same sentence โ not because it has more words, but because of how tokenizers are trained. Here is the actual size of the gap and three ways to shrink your bill.
Updated August 12, 2026 ยท LW Forge
The three Claude tiers are priced 5:2:1 for identical traffic, so the choice multiplies your whole bill by a constant. Here's the price table, a worked monthly example, and the use cases where each tier earns its rate.
Updated August 12, 2026 ยท LW Forge
Sonnet costs exactly 2x Haiku on identical traffic. Here's a worked pipeline where that gap is $1,100 a month, the tasks where Haiku matches Sonnet in practice, and a routing pattern that captures most of the savings without the quality risk.
Updated August 4, 2026 ยท LW Forge
ChatGPT, Claude and Gemini all have a free tier, and all three ration it differently โ by messages, by a rolling time window, or by compute. Here is what a student or self-learner can realistically run on free access in 2026, and when it stops being enough.
Updated August 12, 2026 ยท LW Forge
GLM-5.2 lists at $1.40/$4.40 direct and about half that through aggregators. A worked comparison against GPT-5.6 and Claude, the cache detail that quietly narrows the gap, and the volume where self-hosting the open weights finally beats the API.
Updated August 12, 2026 ยท LW Forge
The same H100 rents for $1.65 or $11.06 an hour depending on where you look. A provider-by-provider table, the spot/reserved discounts that actually apply, and the breakeven math that decides whether self-hosting beats an API bill.
Updated August 12, 2026 ยท LW Forge
Moonshot's 2.8T open-weights flagship lists at $3/$15 โ 40% under Opus 4.8, and, since Anthropic made Sonnet 5's launch price permanent, 50% above it too. A worked comparison, what the 1M context window is actually worth against Claude's 200k, and where the premium is justified.
Updated August 4, 2026 ยท LW Forge
Nemotron 3 Ultra lists at $0.50/$2.20 per million tokens and has a genuinely free tier. What that free tier costs you in data terms, what the same model costs to self-host on eight GPUs, and why open weights and a cheap API solve different problems.
Updated August 4, 2026 ยท LW Forge
OpenAI bills in dollars, but the invoice lands in reais with IOF, a card spread and โ for companies โ a stack of import taxes on top. The real multiplier is not the exchange rate, and it is closer to 1.5x than to 1.05x.
Updated August 12, 2026 ยท LW Forge
A worked breakdown of what one RAG query costs โ embedding, retrieved context and generation โ on 10,000 queries a month. The embedding is a rounding error, the answer length is not, and the re-embedding bill is the one nobody budgets for.
Updated August 3, 2026 ยท LW Forge
128k was the ceiling not long ago; 1M is standard on several models now. What that translates to in pages and books, what it costs to fill, and why bigger isn't automatically better.
Updated August 3, 2026 ยท LW Forge
A 200k window holds about 400 dense pages, a million-token window about 2,000 โ the real number for your document depends on density, and here's how to check it.
Updated July 27, 2026 ยท LW Forge
The token math behind PDF summarization, whether a 300-page document even fits in one request, and when the free chat apps stop being enough.
Updated July 27, 2026 ยท LW Forge
How OpenAI's automatic caching and Anthropic's explicit breakpoints actually price out, a worked example showing the savings, and the three situations where caching quietly does nothing.
Updated July 20, 2026 ยท LW Forge
The budget tiers of OpenAI, Anthropic, Google and DeepSeek compared honestly: list prices, what 'cheap per token' hides, and how to pick the cheapest model that's still correct.
Updated July 20, 2026 ยท LW Forge
DeepSeek's V4 lineup undercuts Western flagships by 10-40x on list price. What you actually trade away โ latency, ecosystem, data governance โ and the workloads where it wins outright.
Updated July 13, 2026 ยท LW Forge
What $20/month buys in API tokens, where the breakeven sits for each GPT-5.6 tier, and a simple rule for choosing between ChatGPT Plus, Claude Pro and pay-per-use.
Updated August 12, 2026 ยท LW Forge
Why price-per-million comparisons mislead, how the Anthropic and OpenAI tiers actually map onto each other, and how to measure cost per task instead of cost per token.
Updated July 6, 2026 ยท LW Forge
Open math for a support bot, an internal assistant and a complex agent on the GPT-5.6 family โ plus the history trap and caching discounts that change the bill.
Updated July 6, 2026 ยท LW Forge
What tokens actually are, why every AI provider bills in them, and how to convert words, characters and pages into tokens without getting your budget wrong.