1. Home
  2. Blog
  3. AI Tools

Gemini 3.8 Flash: Price, Benchmarks, and How to Use It Free

Lindy Drope
Lindy Drope
Founding GTM at Lindy
Lindy leads GTM at Lindy and is the team’s most prolific automation builder. She publishes weekly educational videos and articles on building AI assistants – And yes, she’s a real person!
Lindy Drope
Written by
Lindy Drope
Flo Crivello
Flo Crivello
Founder and CEO of Lindy
Flo Crivello is the founder and CEO of Lindy. Before that, he founded Teamflow and was a product manager at Uber. He writes about technology, startups, and the future of work on his blog.
Flo Crivello
Reviewed by
Flo Crivello
Last Updated:
October 1, 2026
Expert Verified

Google has shipped three Flash models in six weeks: 3.6 Flash in July, 3.7 Flash in August, and Gemini 3.8 Flash on September 2. (If you just finished swapping your model string to 3.7, I'm sorry.)

I went through Google's launch post, the Gemini API pricing page and release notes, the model card, the rate-limit docs, and the Google AI plans page, so every price and score below traces back to its source, and almost all of them to Google's own pages.

Here's what Gemini 3.8 Flash is, what it costs now and after the price doubles in January, how to use it without paying, how it scores against Claude and GPT, and why it can feel slow.

TL;DR:

  • What it is: Google's newest Flash model, which Google calls its "most intelligent workhorse model," built on Gemini 3.7 Flash.
  • Released: September 2, 2026, generally available in the Gemini API on day one.
  • API price per million tokens: $0.75 input and $3.75 output through December 31, 2026, then $1.50 and $7.50 from January 1, 2027.
  • Is it free? Yes on the Gemini API free tier and in Google AI Studio. In the Gemini app, you need Google AI Pro or Ultra.
  • Biggest caveat: it spends more tokens and time on hard tasks by design, so turn the thinking level down when you don't need the extra effort.

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's newest Flash model, released on September 2, 2026, and available to developers, businesses, and paying Gemini subscribers the same day. Google's launch post calls it "our best reasoning and coding model yet, at the same speed and low cost of 3.7."

Flash is the fast, lower-cost tier of the Gemini family. It's the model you pick when you want strong results on a lot of requests without paying top-tier prices for each one.

The model card says 3.8 Flash is based on 3.7 Flash, so it's a tuned-up successor to the model you may already use. Google has kept up a steady stream of point releases since the Gemini 3 generation arrived last November.

The specs, from Google's model page and model card:

🔧 Spec 📋 Gemini 3.8 Flash
API model ID gemini-3.8-flash
Context window 1,048,576 input tokens
Max output 65,536 tokens
Inputs Text, image, video, audio, PDF
Output Text only
Thinking levels Low, medium (default), high
Knowledge cutoff March 2026
Computer use Supported (preview)

One small catch on that cutoff: the model card says that in some domains its knowledge may stop at January 2025, in line with the rest of the Gemini 3 family. For anything recent, turn on grounding with Google Search (a paid-tier feature).

What is Gemini 3.8 Flash Cyber?

Gemini 3.8 Flash Cyber is a security version of 3.8 Flash that Google launched the same day, built on the same foundation and tuned for finding and patching software vulnerabilities. Google calls it "our most capable cybersecurity model."

You probably can't use it, though (and that's on purpose). Because it ships with looser safety limits for security work, Google offers it only to vetted defenders, like government agencies, critical infrastructure operators, and software maintainers, through Google DeepMind's new Fairwind Program.

Google's early numbers are strong. It says its Chrome security team got 2.6 times more correct patches with it than with the best, much larger commercial models, and on the CWE-Bench test it scored 47.2%, a hair behind Anthropic's Claude Fable 5 at 47.8%.

Gemini 3.8 Flash pricing

Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on the paid Gemini API tier, but that's an introductory price. On January 1, 2027, it doubles to $1.50 and $7.50.

These are the prices from Google's Gemini API pricing page, per million tokens, as of September 30, 2026:

💳 Tier 📥 Input (now → 2027) 📤 Output (now → 2027) ♻️ Cached input (now → 2027)
Standard $0.75 → $1.50 $3.75 → $7.50 $0.075 → $0.15
Batch $0.375 → $0.75 $1.875 → $3.75 $0.0375 → $0.075
Flex $0.375 → $0.75 $1.875 → $3.75 $0.0375 → $0.075
Priority $1.35 → $2.70 $6.75 → $13.50 $0.135 → $0.27

A few pricing rules worth knowing before you budget:

  1. Thinking tokens count as output. The output price includes the tokens the model spends reasoning, and 3.8 Flash spends a lot of them on hard tasks (more on that below).
  2. Batch and Flex are half price. Use either one for any job that can wait.
  3. There's no long-prompt surcharge. Google lists one price for the whole 1-million-token window.
  4. Search grounding has its own meter. On the paid tier, you get 5,000 search requests a month free, shared across Gemini 3 models, then it's $14 per 1,000. One prompt can trigger several searches, and each one counts.
  5. 3.7 Flash and 3.6 Flash get the same discount. Google's introductory pricing covers all three Flash models through December 31, 2026, according to Google's developer docs.

What that looks like on a real bill

Say your app sends 1,000 requests, each with 5,000 input tokens and 2,000 output tokens. On standard pricing today, that's $3.75 for input plus $7.50 for output, so $11.25 total.

From January 1, the same job costs $22.50. Run it through the Batch API and it drops back to $11.25.

For comparison, here are the prices Google printed next to 3.8 Flash in its own benchmark table:

🤖 Model 📥 Input per 1M 📤 Output per 1M
Gemini 3.8 Flash $0.75 ($1.50 from 2027) $3.75 ($7.50 from 2027)
Claude Opus 5 $5.00 $25.00
Claude Sonnet 5 $2.00 $10.00
GPT-5.6 Sol $4.00 $20.00
GPT-5.6 Terra $2.00 $12.00

Even at the 2027 price, 3.8 Flash is the cheapest model in that lineup per token. A model that burns more tokens per task can close that gap, so test your real workload before you trust the table.

Is Gemini 3.8 Flash free?

Yes, Gemini 3.8 Flash is free on the Gemini API's free tier and in Google AI Studio, with a few strings attached. In the consumer Gemini app, you'll need a paid plan.

Here's how to use it for free:

  1. Google AI Studio. Sign in at aistudio.google.com with a Google account, pick Gemini 3.8 Flash, and start prompting. It's the fastest way to try it, with no code and no billing account.
  2. The Gemini API free tier. Create an API key in AI Studio and call gemini-3.8-flash. Standard input and output tokens are free of charge on the free tier.

The strings attached:

  • Limits aren't published. Google's rate-limits page says free-tier limits are shown in AI Studio, and it doesn't guarantee them.
  • Some features are paid-only. Batch, Flex, and grounding with Google Search aren't available on the free tier.
  • Google can use what you send. The pricing page says free-tier content is used to improve Google's products, while paid-tier content isn't. Keep sensitive or client data off the free tier.

In the Gemini app, you need Google AI Pro or Ultra. Google says 3.8 Flash is for Google AI Pro and Ultra subscribers in the Gemini app, AI Mode in Search, and Gemini in Google Sheets.

Google's plans page still lists 3.6 Flash for the free app, and the $4.99-a-month AI Plus plan isn't one of the plans Google named for 3.8 Flash.

That means the cheapest way to get it in the app is Google AI Pro at $19.99 a month. Ultra starts at $99.99 a month.

Gemini 3.8 Flash benchmarks

Gemini 3.8 Flash beats 3.7 Flash on every test in Google's table, and it wins several outright against much pricier models, though Claude Opus 5 still tops a few. These are Google's own numbers from the launch post and model card, so read them as the vendor's best case:

📊 Benchmark ⚡ 3.8 Flash ⏮️ 3.7 Flash 🟠 Claude Opus 5 🟢 GPT-5.6 Sol
DeepSWE v1.1 (coding) 73.7% 65.3% 74.0% 72.7%
Terminal-bench 2.1 (terminal coding) 89.4% 85.8% 89.1% 88.8%
Terminal-bench 4.0 (agent tasks) 19.1% 11.2% 51.8% 37.3%
OSWorld-2.0 (computer use) 59.0% 50.6% 75.4% 62.6%
Vals Finance Agent v2 61.4% 59.0% 58.6% 53.8%
Harvey's Legal Agent Benchmark 10.0% 8.8% 6.7% 2.5%
HLE-Verified (expert reasoning) 54.9% 53.6% 54.4% 54.5%
CharXiv Reasoning (charts) 86.2% 84.5% 83.7% 85.8%
LVBench (long video) 87.8% (agentic) 85.4% 75.4% 82.1%

Google's full table also includes Claude Sonnet 5 and GPT-5.6 Terra, plus knowledge-work and biology tests. Here's how I'd read it:

  • Coding made a big leap. DeepSWE v1.1 went from 65.3% to 73.7%, which puts a Flash model within 0.3 points of Claude Opus 5 and a point ahead of GPT-5.6 Sol.
  • Finance and legal work are its best wins. It posts the top score on Vals Finance Agent v2 and Harvey's legal test, though every model scores low on Harvey's, where 10.0% is the best result in the table.
  • HLE-Verified is a photo finish. The 54.9% score that Google highlights beats GPT-5.6 Sol by 0.4 points and Claude Opus 5 by 0.5.
  • Long agent runs and computer use still belong to Opus 5. Claude Opus 5 scores 51.8% on Terminal-bench 4.0 against 19.1% for 3.8 Flash, and 75.4% on OSWorld-2.0 against 59.0%.
  • Video is a real strength. On LVBench, a long-video test, it beats Opus 5 by more than 12 points.

The short version: at today's introductory price, 3.8 Flash costs roughly a seventh of what Opus 5 does per token, and it lands shockingly close on coding and reasoning. It trails most on long, multi-step agent work.

If you're choosing a model for coding work, our roundup of AI coding agents covers the AI coding assistants worth testing.

OpenAI also put Gemini 3.8 Flash in its own launch table for GPT-6 Astra, where it scored 95.3% on GPQA Diamond. We break down that table in our GPT-6 explainer.

Why Gemini 3.8 Flash feels slow (and how to speed it up)

Gemini 3.8 Flash can feel slower than 3.7 Flash because Google built it to work harder on difficult tasks. Google's launch post puts it plainly: "3.8 Flash works harder." On complex work, it takes extra reasoning steps and calls tools again and again before it answers.

Google also says the model "might use more tokens to maximize performance, especially at higher effort levels." The model card goes further and warns about "occasional slowness or timeout issues."

That extra work is why some requests can sit for a while in AI Studio before an answer shows up. It's also why your bill can climb faster than the per-token price suggests, since all that extra thinking is billed as output.

The fix is the thinking level. Gemini 3.8 Flash has three settings, which you pick with the thinking_level parameter in the API:

🎚️ Thinking level 🛠️ What Google says it does 🎯 Good for
Low Cuts time-to-answer for latency-critical tasks Chat, extraction, tagging, short answers
Medium (default) Best quality for most tasks Most work, including complex code and agents
High Maximizes reasoning and tool use Deep reasoning, math, hard multi-step jobs

A few tips that save time and money:

  • Start low for simple jobs. Google's developer docs say that for everyday tasks, you can lower the reasoning effort to reduce token use.
  • Skip the "minimal" setting. Google says minimal isn't supported on 3.8 Flash and will return an error (so if your old code sets it, that's your bug).
  • Update old code. Google's migration notes say to replace thinking_budget with thinking_level and strip temperature, top_p, and top_k from your generation settings.
  • Keep 3.7 Flash as a fallback. Google says 3.7 Flash "remains fully supported," and its deprecations page lists no shutdown date, so you can route quick, simple requests there.

Where you can use Gemini 3.8 Flash

You can use Gemini 3.8 Flash in Google's consumer apps, its developer tools, and a handful of third-party platforms. Where you get it depends on whether you're chatting, building, or deploying at work.

For everyday use

  1. The Gemini app. Available to Google AI Pro and Ultra subscribers. If you've been weighing Gemini against ChatGPT and other chatbots, our guide to AI tools like ChatGPT compares the main options.
  2. AI Mode in Google Search. Also for AI Pro and Ultra subscribers.
  3. Gemini in Google Sheets. Same plans, and a handy spot for a fast model that's strong on finance tasks.

For developers

  1. Google AI Studio. The free, browser-based playground for testing prompts and getting an API key.
  2. The Gemini API. Call it as gemini-3.8-flash on the free or paid tier.
  3. Google Antigravity. Google's platform for building with coding agents. The Antigravity agent in Gemini Managed Agents and the Antigravity SDK both use 3.8 Flash by default, according to Google's docs.
  4. Android Studio. Google lists it alongside AI Studio as a way for Android app developers to reach 3.8 Flash through the Gemini API.
  5. Stitch. Google's tool for generating app interfaces from a prompt.

For businesses

  1. Gemini Enterprise. Companies get 3.8 Flash through Gemini Enterprise Agent Platform (the product formerly called Vertex AI), which adds regional hosting, provisioned throughput, and volume discounts.

Through third-party platforms

  1. Cloudflare. Cloudflare's model page lists it at Google's prices, with the full 1-million-token context window.
  2. OpenRouter. OpenRouter lists google/gemini-3.8-flash at $0.75 and $3.75, plus a half-price batch version.

Before you paste in a giant codebase, check the context limit in whichever tool you're using.

‍

{{templates}}

‍

The rest of Gemini 3.8

Gemini 3.8 is also a family of voice and speech models that Google shipped in the weeks after 3.8 Flash. They share the version number, but they're separate models built for audio, and 3.8 Flash itself doesn't support the Live API.

🎙️ Model 📅 Released 🎯 What it does 📍 Where to use it
Gemini 3.8 Live Sept 15 Low-latency voice conversations in 97 languages Gemini API, AI Studio, Search Live
Gemini 3.8 Live Extended Thinking Sept 15 Voice conversations with background reasoning Gemini API, AI Studio, Gemini Live, Docs, Gmail, Keep
Gemini 3.8 Live with Live Avatar Sept 24 Real-time video avatar with lip-sync Gemini Enterprise
Gemini 3.8 Flash TTS Sept 22 Expressive, creative text-to-speech Gemini API, AI Studio, Gemini Notebook
Gemini 3.8 Flash-Lite TTS Sept 22 Faster, cheaper text-to-speech Gemini API, AI Studio, Google Vids

Google's Live launch post says the Extended Thinking model reaches Docs on AI Pro and Ultra, and Gmail and Keep for all Google AI subscribers.

The two speech models also have introductory prices on the pricing page. Flash TTS costs $9 per million audio output tokens (about $0.00225 per 10 seconds of speech), and Flash-Lite TTS costs $6, with both doubling in 2027.

Should you switch from Gemini 3.7 Flash?

Switch for hard work and keep 3.7 Flash around for the easy stuff. Since both cost the same per token, now and after January, the trade-off is speed and token use against quality.

Switch to Gemini 3.8 Flash if you:

  • Run coding, finance, legal, or long-video tasks, where its scores are strongest
  • Build agents that call tools, since it works in smaller steps and checks its work along the way
  • Want a Flash-priced model that gets within a point of Claude Opus 5 on DeepSWE

Stay on Gemini 3.7 Flash if you:

  • Need the fastest possible answers for chat, classification, or simple extraction
  • Run high-volume jobs where extra thinking tokens would pile up on the bill
  • Have prompts tuned for 3.7 that already work well

Look at other models if you:

  • Run long, multi-step agent work or computer-use tasks, where Claude Opus 5 leads Google's own table by a wide margin
  • Want open weights you can run yourself, where Qwen 3.8 is worth testing
  • Want a wider shortlist of AI tools worth trying, starting with the ones I use most

Test Gemini 3.8 Flash for a week before the price doubles

The best thing about the introductory price is that it gives you three months to find out whether Gemini 3.8 Flash earns its spot, before the per-token price doubles on January 1.

A simple one-week test:

  1. Pick 20 real tasks. Pull them from your own workload, including a few hard ones and a few boring ones.
  2. Run each one on 3.7 Flash and 3.8 Flash. Use low, medium, and high thinking on 3.8, and note the answer quality, the time to answer, and the tokens used.
  3. Price it at 2027 rates. Multiply your token counts by $1.50 and $7.50 per million, since that's what you'll pay once the discount ends.

If 3.8 wins on the hard tasks and ties on the easy ones, route the hard work to 3.8 and leave the rest on 3.7. If it only wins at high thinking, check whether the better answers are worth the extra tokens.

The same test works inside other tools, too. Lindy, the AI teammate that works in Slack, lets you pick the model for each task, so if your team uses it, check whether Gemini 3.8 Flash is on the model list before you run the test there.

‍

{{cta}}

‍

FAQ

Is Gemini 3.8 Flash free?

Yes, Gemini 3.8 Flash is free in Google AI Studio and on the Gemini API free tier, within limits Google shows in AI Studio. Free-tier data is used to improve Google's products. In the Gemini app, it needs a Google AI Pro ($19.99 a month) or Ultra plan.

How much does Gemini 3.8 Flash cost?

Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens on the paid API through December 31, 2026. From January 1, 2027, it costs $1.50 and $7.50. Batch and Flex requests cost half.

Is Gemini 3.8 Flash better than 3.7 Flash?

Yes, Gemini 3.8 Flash beats 3.7 Flash on every benchmark in Google's table, including 73.7% vs 65.3% on DeepSWE v1.1 and 19.1% vs 11.2% on Terminal-bench 4.0.

It uses more tokens on hard tasks, so 3.7 Flash can still be the faster pick for simple jobs, and cheaper per task because it spends fewer tokens.

Why is Gemini 3.8 Flash so slow?

Gemini 3.8 Flash is slow on some tasks because it takes extra reasoning steps and calls tools repeatedly to get better answers. Set the thinking level to low for quick tasks, keep medium for most work, and save high for hard problems.

Is Gemini 3.8 Flash good?

Yes, Gemini 3.8 Flash is one of the strongest low-cost models available, with top scores in Google's table on finance, legal, long-video, and HLE-Verified tests. Claude Opus 5 still leads it on long agent tasks and computer use.

Is Gemini 3.8 Flash better than ChatGPT?

Yes, on several of Google's benchmarks. Gemini 3.8 Flash beats GPT-5.6 Sol on DeepSWE v1.1, Vals Finance Agent v2, and HLE-Verified, at about a fifth of the per-token price during the introductory period.

GPT-5.6 Sol wins on Terminal-bench 4.0 and computer use, and OpenAI's newer GPT-6 models weren't in Google's comparison.

What is Gemini 3.8 Flash best for?

Gemini 3.8 Flash is best for coding, finance and legal agent tasks, long-video analysis, and high-volume work where cost per request matters. It leads Google's table on Vals Finance Agent v2, Harvey's Legal Agent Benchmark, and LVBench.

Is Gemini 3.8 Flash better than Gemini 3.1 Pro?

Gemini 3.8 Flash is newer and much cheaper than Gemini 3.1 Pro, and Google calls 3.8 its "best reasoning and coding model yet," though it hasn't published a head-to-head benchmark between the two.

Gemini 3.1 Pro is still a preview model, last updated in February 2026, and costs $2 per million input tokens and $12 per million output tokens with no free tier. Start with 3.8 Flash and save 3.1 Pro for a test on your hardest jobs.

Can you download Gemini 3.8 Flash?

No, you can't download Gemini 3.8 Flash. It only runs as a hosted model, through Google's own apps and APIs (the Gemini app, the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, AI Mode, and Antigravity) and through hosts like Cloudflare and OpenRouter.

Google also doesn't publish its parameter count.

Save 2 Hours Every Day
Lindy is your ultimate AI assistant that manages inbox, meetings, and follow-ups—so you stay ahead of the chaos.
Try Lindy for Free
About the editorial team
Lindy Drope
Lindy Drope
Founding GTM at Lindy

Lindy leads GTM at Lindy and is the team’s most prolific automation builder. She publishes weekly educational videos and articles on building AI assistants – And yes, she’s a real person!

Flo Crivello
Flo Crivello
Founder and CEO of Lindy

Flo Crivello is the founder and CEO of Lindy. Before that, he founded Teamflow and was a product manager at Uber. He writes about technology, startups, and the future of work on his blog.

Ready when you are.

Free to try. In your Slack in two minutes.

Try for free
7-day free trial • Cancel anytime