All articles

Gemini 3.7 Flash Launches at Half the Price of 3.5 — But Only Until 2027

By LW Forge — maintainer of LLM Scout · Updated August 31, 2026

Google shipped Gemini 3.7 Flash in mid-August 2026, and the headline isn't the model's capabilities — it's the price. Confirmed directly on Google's own Gemini API pricing page, the new Flash tier launches at roughly half of what Gemini 3.5 Flash costs today, with one catch worth planning around: the rate is explicitly temporary.

The actual numbers

ModelInput ($/1M tokens)Output ($/1M tokens)Notes
Gemini 3.5 Flash$1.50$9.00Current-gen Flash, flat rate
Gemini 3.7 Flash (through Dec 31, 2026)$0.75$3.75Introductory rate
Gemini 3.7 Flash (from Jan 1, 2027)$1.50$7.50Standard rate takes over

Input drops 50%, and output drops 58% for the rest of 2026. From January 1, 2027, the input rate returns to parity with 3.5 Flash ($1.50) while output settles at $7.50 — still 17% cheaper than 3.5 Flash's $9.00, just not the introductory discount. Google's own pricing table lists both rates side by side, so this isn't a rumor or a limited-time promo code — it's the published schedule for the model.

This is a different shape of "temporary" than the price cuts we've tracked before. When OpenAI cut GPT-5.6 Luna and Terra in July 2026, that reduction was permanent — no expiration date. When Claude Sonnet 5's launch pricing became permanent in August, that removed an expiration that had originally been scheduled. Gemini 3.7 Flash runs the opposite direction: a real discount with a real sunset date already on the books. If you're budgeting a product launch or a multi-month contract around this rate, build the January 1, 2027 step-up into your model, not just today's number.

Where it lands versus the rest of the cheap tier

Gemini 3.7 Flash's introductory rate is competitive but not the outright cheapest option on the market. For context, at time of writing: OpenAI's GPT-5.6 Luna sits at $0.20/$1.20, GLM-5.2 at $1.40/$4.40, and Claude Haiku 4.5 at $1.00/$5.00. Gemini 3.7 Flash's $0.75/$3.75 introductory rate undercuts Haiku and GLM-5.2 on both input and output, but Luna remains meaningfully cheaper on a pure per-token basis. Where Gemini 3.7 Flash tends to win is the combination of a large context window with Google's broader tooling (grounding, function calling, native multimodal input) at a price closer to the budget tier than to the mid tier — see the full rundown of where each provider sits in the cheapest LLM APIs of 2026.

What it costs in practice

Run the numbers on a moderate workload: 2,000 input tokens and 500 output tokens per request, 50,000 requests a month (a common shape for a support-triage or summarization feature).

  • At today's 3.5 Flash rate: (2,000 × $1.50 + 500 × $9.00) / 1,000,000 × 50,000 = $375 in input + $225 in output = $600/month.
  • At 3.7 Flash's introductory rate: (2,000 × $0.75 + 500 × $3.75) / 1,000,000 × 50,000 = $75 in input + $93.75 in output = $168.75/month.
  • At 3.7 Flash's 2027 standard rate: (2,000 × $1.50 + 500 × $7.50) / 1,000,000 × 50,000 = $150 + $187.50 = $337.50/month.

That's a 72% reduction through the end of 2026, settling to a still-real 44% reduction once the standard rate takes over — on the same workload, just by migrating from 3.5 Flash to 3.7 Flash. Run your own request volume through the token calculator to see where your workload lands, and use the GPT vs Claude cost comparison if Gemini is one of several providers you're weighing rather than a default choice.

Why the expiration date matters more than usual

Most price cuts in this market land as flat, unconditional reductions — OpenAI's July cut to Luna and Terra didn't come with fine print about reverting. Gemini 3.7 Flash's structure is closer to an interest rate: real for now, scheduled to reset. That distinction changes how you should treat it in planning, not just in this month's invoice.

Two groups should pay closer attention than a headline price comparison usually warrants. Teams pricing a subscription product around Gemini 3.7 Flash's introductory rate need to model the January 2027 step-up into their margin, not just today's number — a 44% jump in output cost per token is enough to turn a comfortable margin into a thin one if it isn't priced in from day one. And teams running short-lived pilots or hackathon-scale projects that will wrap up before year-end can treat the introductory rate as the real, durable number for the life of the project — there's no risk of the rug being pulled mid-pilot, since Google has committed to the schedule on its own pricing page rather than leaving it as a silent, revocable promotion.

The practical takeaway

If you're already on Gemini 3.5 Flash, migrating to 3.7 Flash is close to a free upgrade — same tier, lower price, and Google is explicitly positioning it as the replacement, the same pattern as 3.5 Flash replacing 2.5 Flash earlier in 2026. If you're choosing a provider for a new project and pure per-token cost is the deciding factor, it's still worth comparing against Luna and GLM-5.2 rather than assuming Flash's headline discount makes it the cheapest option outright — it doesn't, quite. What it does do is narrow the gap, and it gives Google's stack a credible budget tier through the rest of 2026 that didn't really exist a month ago.

Frequently asked questions

When does the Gemini 3.7 Flash introductory price end?

The $0.75/$3.75 per million token rate holds through December 31, 2026. Starting January 1, 2027, the standard rate of $1.50/$7.50 takes over, per Google's own pricing page.

Is Gemini 3.7 Flash cheaper than Gemini 3.5 Flash?

Yes, on both the introductory and the 2027 standard rate. Introductory pricing cuts input 50% and output 58% versus 3.5 Flash; the 2027 standard rate still keeps output 17% cheaper, with input at parity.

Is Gemini 3.7 Flash the cheapest LLM API available?

No. OpenAI's GPT-5.6 Luna, at $0.20/$1.20 per million tokens, remains cheaper on a pure per-token basis. Gemini 3.7 Flash's advantage is combining a large context window and Google's broader tooling at a price closer to the budget tier.

Pricing verified directly against Google's Gemini API pricing page on 2026-08-29.