Gemini 3.8 Flash Hits Stable GA at $0.75 per Million In
Flash tiers don't win benchmark charts. They win invoices — and this one lands thirteen times under the frontier rate.
Jump to 5 sections
Google moved Gemini 3.8 Flash to stable general availability on September 2, 2026, one day after Anthropic's Fable 5.1 and one day before OpenAI's GPT-6 Astra. It is the cheapest thing in that week's release cluster by a wide margin, and for a lot of production traffic that makes it the most consequential.
Flash tiers do not get launch-day attention because they do not win benchmark charts. They win invoices.
The price is the product
Google listed Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output, with that rate stated as holding through December 31, 2026. The end date matters — promotional pricing with a stated expiry is a budget line that changes on a known day.
Set that against the frontier tier and the gap is not subtle. A flagship at $10 input and $50 output costs roughly thirteen times more per token in both directions.
| Tier | Input / 1M | Output / 1M | Cost per 10k runs* |
|---|---|---|---|
| Frontier flagship | ~$10 | ~$50 | $20,000 |
| Gemini 3.1 Pro | $2 | $12 | $4,800 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $1,500 |
*40,000 output tokens per run, output only. Verify current rates against Google's published Gemini API pricing before budgeting, because promotional windows move.
This lands in the middle of the densest release week of the year, alongside OpenAI's GPT-6 Astra at the opposite end of the price ladder. Two launches, two days apart, aimed at almost entirely different buyers.
The routing math here is the whole argument. Most production requests — classification, extraction, routing, summarization of things nobody will litigate — do not produce a measurably different outcome on a flagship. They just cost thirteen times more. We work through that ladder in our breakdown of LLM API pricing across the 2026 tiers.
What stable GA actually promises
Stable GA is a support commitment more than a capability statement. It means the model identifier will not change under you, the behavior is not expected to drift mid-quarter, and there is a deprecation process rather than an abrupt cutoff.
For anyone running scheduled jobs against a pinned model string, that is the only thing on the announcement that matters. A preview endpoint can change tokenization behavior, output formatting, or default parameters between Tuesday and Thursday. A stable one is not supposed to.
Read the deprecation policy, not the launch post. Google publishes model lifecycle and versioning terms in its Gemini API model documentation, and that page tells you how much notice you get when a version retires. Launch announcements essentially never carry that information, which is one of the omissions our guide to reading launch announcements flags.
Latency is the other half of a Flash tier
Cheap and fast usually travel together, and for user-facing work the latency side often justifies the tier on its own. A model that returns first token in a few hundred milliseconds changes what interactions are possible; one that takes four seconds does not, regardless of quality.
Google did not publish comparative latency figures in the GA announcement, which is normal and mildly annoying. Measure it yourself on your own payloads and your own region — published throughput numbers rarely survive contact with a real prompt distribution.
Three things to measure before you commit: time to first token under realistic input length, tokens per second at your p50 and p99, and how both degrade when you enable tool calling. Tool-calling round trips frequently dominate end-to-end latency in ways a raw generation benchmark never shows.
Where Flash tiers stop making sense
Long multi-step agent chains are the honest limit. Error rates compound across steps, so a small per-step quality gap becomes a large end-to-end failure rate, and the cheaper model stops being cheaper once you count retries and cleanup.
At Model Drop our working rule is that a cheap tier is right until a wrong answer costs more than the inference does. That threshold arrives sooner in agentic workloads than most teams expect, and it is worth measuring rather than assuming in either direction.
The sensible architecture is a router, not a standardization. Send the bulk of traffic to Flash, escalate the requests that fail a confidence or validation check to a higher tier, and measure what fraction actually escalates. For most teams that fraction is small enough to make the frontier tier a rounding error on the bill.
What we are watching
Whether the December 31 promotional rate holds or resets, and whether independent latency measurements match the positioning. A Flash tier that is cheap but slow is a different product from one that is cheap and fast, and only the second one changes architectures.
By Derek Plummer, Staff Writer at Model Drop. Reported September 2, 2026. Pricing figures are Google's published list rates and are subject to stated promotional end dates.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.