Google has shipped three Flash models in six weeks: 3.6 Flash in July, 3.7 Flash in August, and Gemini 3.8 Flash on September 2. (If you just finished swapping your model string to 3.7, I'm sorry.)
I went through Google's launch post, the Gemini API pricing page and release notes, the model card, the rate-limit docs, and the Google AI plans page, so every price and score below traces back to its source, and almost all of them to Google's own pages.
Here's what Gemini 3.8 Flash is, what it costs now and after the price doubles in January, how to use it without paying, how it scores against Claude and GPT, and why it can feel slow.
TL;DR:
Gemini 3.8 Flash is Google's newest Flash model, released on September 2, 2026, and available to developers, businesses, and paying Gemini subscribers the same day. Google's launch post calls it "our best reasoning and coding model yet, at the same speed and low cost of 3.7."
Flash is the fast, lower-cost tier of the Gemini family. It's the model you pick when you want strong results on a lot of requests without paying top-tier prices for each one.
The model card says 3.8 Flash is based on 3.7 Flash, so it's a tuned-up successor to the model you may already use. Google has kept up a steady stream of point releases since the Gemini 3 generation arrived last November.
The specs, from Google's model page and model card:
One small catch on that cutoff: the model card says that in some domains its knowledge may stop at January 2025, in line with the rest of the Gemini 3 family. For anything recent, turn on grounding with Google Search (a paid-tier feature).
Gemini 3.8 Flash Cyber is a security version of 3.8 Flash that Google launched the same day, built on the same foundation and tuned for finding and patching software vulnerabilities. Google calls it "our most capable cybersecurity model."
You probably can't use it, though (and that's on purpose). Because it ships with looser safety limits for security work, Google offers it only to vetted defenders, like government agencies, critical infrastructure operators, and software maintainers, through Google DeepMind's new Fairwind Program.
Google's early numbers are strong. It says its Chrome security team got 2.6 times more correct patches with it than with the best, much larger commercial models, and on the CWE-Bench test it scored 47.2%, a hair behind Anthropic's Claude Fable 5 at 47.8%.
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on the paid Gemini API tier, but that's an introductory price. On January 1, 2027, it doubles to $1.50 and $7.50.
These are the prices from Google's Gemini API pricing page, per million tokens, as of September 30, 2026:
A few pricing rules worth knowing before you budget:
Say your app sends 1,000 requests, each with 5,000 input tokens and 2,000 output tokens. On standard pricing today, that's $3.75 for input plus $7.50 for output, so $11.25 total.
From January 1, the same job costs $22.50. Run it through the Batch API and it drops back to $11.25.
For comparison, here are the prices Google printed next to 3.8 Flash in its own benchmark table:
Even at the 2027 price, 3.8 Flash is the cheapest model in that lineup per token. A model that burns more tokens per task can close that gap, so test your real workload before you trust the table.
Yes, Gemini 3.8 Flash is free on the Gemini API's free tier and in Google AI Studio, with a few strings attached. In the consumer Gemini app, you'll need a paid plan.
Here's how to use it for free:

The strings attached:
In the Gemini app, you need Google AI Pro or Ultra. Google says 3.8 Flash is for Google AI Pro and Ultra subscribers in the Gemini app, AI Mode in Search, and Gemini in Google Sheets.
Google's plans page still lists 3.6 Flash for the free app, and the $4.99-a-month AI Plus plan isn't one of the plans Google named for 3.8 Flash.
That means the cheapest way to get it in the app is Google AI Pro at $19.99 a month. Ultra starts at $99.99 a month.
Gemini 3.8 Flash beats 3.7 Flash on every test in Google's table, and it wins several outright against much pricier models, though Claude Opus 5 still tops a few. These are Google's own numbers from the launch post and model card, so read them as the vendor's best case:
Google's full table also includes Claude Sonnet 5 and GPT-5.6 Terra, plus knowledge-work and biology tests. Here's how I'd read it:
The short version: at today's introductory price, 3.8 Flash costs roughly a seventh of what Opus 5 does per token, and it lands shockingly close on coding and reasoning. It trails most on long, multi-step agent work.
If you're choosing a model for coding work, our roundup of AI coding agents covers the AI coding assistants worth testing.
OpenAI also put Gemini 3.8 Flash in its own launch table for GPT-6 Astra, where it scored 95.3% on GPQA Diamond. We break down that table in our GPT-6 explainer.
Gemini 3.8 Flash can feel slower than 3.7 Flash because Google built it to work harder on difficult tasks. Google's launch post puts it plainly: "3.8 Flash works harder." On complex work, it takes extra reasoning steps and calls tools again and again before it answers.
Google also says the model "might use more tokens to maximize performance, especially at higher effort levels." The model card goes further and warns about "occasional slowness or timeout issues."
That extra work is why some requests can sit for a while in AI Studio before an answer shows up. It's also why your bill can climb faster than the per-token price suggests, since all that extra thinking is billed as output.
The fix is the thinking level. Gemini 3.8 Flash has three settings, which you pick with the thinking_level parameter in the API:
A few tips that save time and money:
You can use Gemini 3.8 Flash in Google's consumer apps, its developer tools, and a handful of third-party platforms. Where you get it depends on whether you're chatting, building, or deploying at work.
Before you paste in a giant codebase, check the context limit in whichever tool you're using.
{{templates}}
Gemini 3.8 is also a family of voice and speech models that Google shipped in the weeks after 3.8 Flash. They share the version number, but they're separate models built for audio, and 3.8 Flash itself doesn't support the Live API.
Google's Live launch post says the Extended Thinking model reaches Docs on AI Pro and Ultra, and Gmail and Keep for all Google AI subscribers.
The two speech models also have introductory prices on the pricing page. Flash TTS costs $9 per million audio output tokens (about $0.00225 per 10 seconds of speech), and Flash-Lite TTS costs $6, with both doubling in 2027.
Switch for hard work and keep 3.7 Flash around for the easy stuff. Since both cost the same per token, now and after January, the trade-off is speed and token use against quality.
Switch to Gemini 3.8 Flash if you:
Stay on Gemini 3.7 Flash if you:
Look at other models if you:
The best thing about the introductory price is that it gives you three months to find out whether Gemini 3.8 Flash earns its spot, before the per-token price doubles on January 1.
A simple one-week test:
If 3.8 wins on the hard tasks and ties on the easy ones, route the hard work to 3.8 and leave the rest on 3.7. If it only wins at high thinking, check whether the better answers are worth the extra tokens.
The same test works inside other tools, too. Lindy, the AI teammate that works in Slack, lets you pick the model for each task, so if your team uses it, check whether Gemini 3.8 Flash is on the model list before you run the test there.
{{cta}}
Yes, Gemini 3.8 Flash is free in Google AI Studio and on the Gemini API free tier, within limits Google shows in AI Studio. Free-tier data is used to improve Google's products. In the Gemini app, it needs a Google AI Pro ($19.99 a month) or Ultra plan.
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on the paid API through December 31, 2026. From January 1, 2027, it costs $1.50 and $7.50. Batch and Flex requests cost half.
Yes, Gemini 3.8 Flash beats 3.7 Flash on every benchmark in Google's table, including 73.7% vs 65.3% on DeepSWE v1.1 and 19.1% vs 11.2% on Terminal-bench 4.0.
It uses more tokens on hard tasks, so 3.7 Flash can still be the faster pick for simple jobs, and cheaper per task because it spends fewer tokens.
Gemini 3.8 Flash is slow on some tasks because it takes extra reasoning steps and calls tools repeatedly to get better answers. Set the thinking level to low for quick tasks, keep medium for most work, and save high for hard problems.
Yes, Gemini 3.8 Flash is one of the strongest low-cost models available, with top scores in Google's table on finance, legal, long-video, and HLE-Verified tests. Claude Opus 5 still leads it on long agent tasks and computer use.
Yes, on several of Google's benchmarks. Gemini 3.8 Flash beats GPT-5.6 Sol on DeepSWE v1.1, Vals Finance Agent v2, and HLE-Verified, at about a fifth of the per-token price during the introductory period.
GPT-5.6 Sol wins on Terminal-bench 4.0 and computer use, and OpenAI's newer GPT-6 models weren't in Google's comparison.
Gemini 3.8 Flash is best for coding, finance and legal agent tasks, long-video analysis, and high-volume work where cost per request matters. It leads Google's table on Vals Finance Agent v2, Harvey's Legal Agent Benchmark, and LVBench.
Gemini 3.8 Flash is newer and much cheaper than Gemini 3.1 Pro, and Google calls 3.8 its "best reasoning and coding model yet," though it hasn't published a head-to-head benchmark between the two.
Gemini 3.1 Pro is still a preview model, last updated in February 2026, and costs $2 per million input tokens and $12 per million output tokens with no free tier. Start with 3.8 Flash and save 3.1 Pro for a test on your hardest jobs.
No, you can't download Gemini 3.8 Flash. It only runs as a hosted model, through Google's own apps and APIs (the Gemini app, the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, AI Mode, and Antigravity) and through hosts like Cloudflare and OpenRouter.
Google also doesn't publish its parameter count.
