Gemini 3.8 Flash beats Claude Opus 5 on two of the benchmarks Google tested it against, and comes within half a point on a third, at roughly one sixth of Opus 5's price (about 15 cents on the dollar). On Harvey's Legal Agent Benchmark it scores a 10.0 percent all-pass rate against Opus 5's 6.7. On Terminal-Bench 2.1, agentic terminal coding, it posts 89.4 percent against Opus 5's 89.1. On DeepSWE v1.1, a long-horizon software engineering benchmark, it lands at 73.7 percent, just under Opus 5's 74.0. All of that against Google's launch pricing for 3.8 Flash: $0.75 per million input tokens and $3.75 per million output, next to $5 and $25 for Opus 5.
The parity is not universal. On Terminal-Bench 4.0, a harder general-agent successor, Opus 5 pulls back ahead decisively: 51.8 percent against 19.1 for 3.8 Flash. Google released Gemini 3.8 Flash on September 2, 2026, its fourth Flash-tier model in under four months, and priced it at Gemini 3.7 Flash's exact rate, with one detail worth flagging before anything else: that launch price is introductory and set to expire.
Where it matches Claude Opus 5, and where it doesn't
The gains over 3.7 Flash are real on paper, and nearly every number is vendor-reported: Google computes most Gemini scores itself and mixes in leaderboard figures for competitors. On HLE-Verified, a hard multidisciplinary reasoning set, 3.8 Flash scores 54.9 percent against 53.6 for 3.7 Flash. On Vals Finance Agent v2, sourced from Vals.AI, it posts 61.4 percent against 59.0. On Harvey's Legal Agent Benchmark it reaches a 10.0 percent all-pass rate against 8.8, and that low ceiling is the quiet headline: Google's own table lists Claude Opus 5 at 6.7 percent and GPT-5.6 Sol at 2.5 on the same test, so complex legal workflows remain mostly unsolved by everyone.
Coding is where the pitch concentrates. On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash scores 73.7 percent, a huge jump from 3.7 Flash at 65.3 and within touching distance of the 74.0 Google's table lists for Claude Opus 5, with GPT-5.6 Sol at 72.7. On Terminal-Bench 2.1, agentic terminal coding, Google's evaluation document lists 89.4 percent, self-computed, against 85.8 for 3.7 Flash, 89.1 for Opus 5, and 88.8 for GPT-5.6 Sol. A 90.8 figure circulated in some early coverage; the number in Google's own document is 89.4.
Two caveats keep the picture honest. First, on Terminal-Bench 4.0, the harder general-agent successor, 3.8 Flash manages 19.1 percent while Opus 5 posts 51.8 in the same table: on genuinely hard agentic work, the budget tier is not close. Second, Google's methodology document carries a correction admitting it originally misreported Opus 5's DeepSWE score because of rounding on a public leaderboard. That is a small, creditable admission, and a useful reminder of how cross-vendor tables get assembled: treat every decimal as directional. Independent measurement lands in the same neighborhood. Artificial Analysis scores 3.8 Flash at 59 on its Intelligence Index with high reasoning and places it on the intelligence versus cost-per-task Pareto frontier.
What Google shipped on September 2
Gemini 3.8 Flash is a text-out workhorse aimed at agentic and coding work, not a multimodal showpiece. The model id is gemini-3.8-flash, the input window holds 1,048,576 tokens, and output caps at 65,536 tokens. It accepts text, images, video, audio, and PDF files, but it only ever writes text back. The knowledge cutoff is March 2026 per the model card. It launched alongside a restricted sibling, Gemini 3.8 Flash Cyber, covered further down, while the frontier tier above Flash stayed unchanged.
- 1,048,576 token input window, 65,536 token output cap, text output only
- Thinking runs at low, medium, or high. The minimal setting returns an error
- Supports context caching, code execution, function calling, structured outputs, Search and Maps grounding, the Batch API, and computer use in preview
- Does not support the Live API, audio generation, or image generation
- Inputs cover text, images, video, audio, and PDFs. Knowledge cutoff is March 2026
Access splits three ways. Developers get it through the Gemini API, Google AI Studio, Android Studio, and Google Antigravity. Enterprises get it through Gemini Enterprise. Consumers get it only behind a paywall: the Gemini app model picker for AI Pro and Ultra subscribers, plus AI Mode in Search and Google Sheets. The most honest line in Google's own launch post is the advice that efficiency-first workloads should stay on Gemini 3.7 Flash, which remains fully supported. That sentence foreshadows the cost numbers below.
The cadence is the subtext. Gemini 3.5 Flash on May 19, 3.6 Flash on July 21, 3.7 Flash on August 13, now 3.8 Flash on September 2: four budget models in under four months, per Google's own release notes, while the frontier models above them, Gemini 3.5 Pro and Gemini 4, stay missing. The Register read the launch as Google reminding everyone it is still in the race, and the reading fits. Fast, cheap, frequently replaced models are how Google is competing right now, and every part of that strategy, the pace, the introductory pricing, the statelessness, eventually shows up somewhere on the bill of anyone who builds on top of it.
Cheap until January 1, then everything doubles
The launch pricing matches Gemini 3.7 Flash to the cent, which makes the upgrade look free. The fine print is the story. Google's pricing page lists every 3.8 Flash rate twice: an introductory figure through December 31, 2026, and a regular figure from January 1, 2027. The doubling is not limited to the headline rates. Batch API requests, cached input reads, and cache storage all double on the same date, and the same schedule applies to Gemini 3.7 Flash and the older 3.6 Flash sitting on the same page. This is not a one-model discount: 3.6, 3.7, and 3.8 Flash all double on the same morning.
| What you pay (per 1M tokens) | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| Input | $0.75 | $1.50 |
| Output | $3.75 | $7.50 |
| Batch API input | $0.375 | $0.75 |
| Batch API output | $1.875 | $3.75 |
| Cached input reads | $0.075 | $0.15 |
| Cache storage, per hour | $0.50 | $1.00 |
Against the models it gets benchmarked beside, even the regular rate stays cheap. At the introductory price, 3.8 Flash runs about 15 cents on the dollar next to Claude Opus 5 at $5 in and $25 out, and about 19 cents next to GPT-5.6 Sol at $4 in and $20 out, using the pricing table in Google's own evaluation document. The January doubling narrows that gap without closing it. The comparison that actually stings is internal, and it shows up per task, not per token.
The catch: per-task cost went up about 45 percent
Identical per-token pricing does not mean identical bills. Artificial Analysis measures cost per task across its Intelligence Index, and Gemini 3.8 Flash lands at $0.58 per task against $0.40 for 3.7 Flash, roughly 45 percent more, driven by about 30 percent more output tokens per task and more turns on agentic evaluations. The model earns its better scores by thinking longer and writing more, and every one of those extra tokens bills at the output rate. This is the same dynamic MemX's earlier breakdown of reasoning token costs documented: the sticker price stays flat while the token count quietly climbs. One lever exists to fight it. Thinking runs at low, medium, or high, and dialing it down trims the reasoning overhead, at some cost to the very scores that justified the upgrade.
How to use Gemini 3.8 Flash free, and what free actually costs
There are exactly two free paths, and both are narrow. The first is not really free: in the consumer Gemini app, the model shows up only for Google AI Pro and Ultra subscribers, so a free account never sees it in the picker. The second is the Gemini API free tier through Google AI Studio, where developers report a limit of 20 requests per day for 3.8 Flash, against 500 per day on the Flash-Lite tier. Google publishes the exact figures inside AI Studio's rate-limit page rather than in its docs, and they can change without a changelog entry, but the launch-week complaint threads are consistent: 20 requests disappears inside a single agent loop once retries and tool calls stack up.
The free tier's real price is data. Google's Gemini API terms state that for unpaid services, Google "uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services," that "human reviewers may read, annotate, and process your API input and output," and, in Google's own words, "Do not submit sensitive, confidential, or personal information to the Unpaid Services." Paid usage flips the deal: Google says it does not use paid-tier prompts or responses to improve its products. Free access to 3.8 Flash exists, but it is 20 requests a day, paid for in training data and subject to human review.
The Cyber variant is not for the public
Gemini 3.8 Flash Cyber, the second model in the announcement, never touches the consumer story. It ships through the Fairwind Program to trusted government authorities, critical-infrastructure operators, and software maintainers. Google's claims for it are specific: a success rate above 70 percent on internal vulnerability-discovery benchmarks spanning 20 programming languages, a 47.2 percent pass@1 on CWE-Bench patching, and 2.6 times more correct patches to Chrome vulnerabilities than the best commercial models. All of that is Google's own reporting, and none of it is available to test from the outside. For everyone reading pricing pages, the Cyber variant is a press release, not a product.
The fourth stateless Flash in four months
Here is what the launch post, the model card, and the developer docs collectively say about memory and personalization: nothing. Four Flash releases in four months, and not one of them ships with any persistence of its own. Whatever Gemini appears to know about a person lives in the app layer: the Gemini app's separate memory and personal-context system, or the Workspace-side personal intelligence setting that governs whether Gemini reads Gmail. The model underneath is a stateless engine that gets swapped every few weeks. Switch surfaces, switch vendors, or watch the app layer change its own rules, and that accumulated context does not come along.
Context caching is the closest thing 3.8 Flash offers to memory, and its pricing structure is honest about what it really is: rent. Cached reads cost $0.075 per million tokens, but keeping tokens cached costs $0.50 per million per hour. Keep a full million-token context warm and that is about $12 a day, roughly $360 a month, and about $720 a month once the January rates land. There is a floor on top of the meter, and it has risen since MemX's earlier analysis of the Gemini caching trap documented it at 1,024 tokens: on 3.8 Flash, the caching docs set the minimum at 4,096 tokens, so shorter prompts never qualify for cache hits at all. A cache is a performance tool that expires by design. Renting your own context back by the hour is a billing model, not a memory.
The churn itself carries a cost that never shows on an invoice. Prompts tuned to one Flash's quirks, few-shot examples that anchored its outputs, guardrails calibrated to its specific failure modes: all of that resets when the next Flash arrives on a three-week cadence. MemX's 103-case stress test of Gemini 3 Flash logged a 27 percent failure rate on real-world memory questions, and every subsequent Flash redraws that failure map without carrying anything forward. Google itself keeps 3.7 Flash fully supported as the efficiency pick, an acknowledgment that the two models behave differently enough for the choice to matter. The only context guaranteed to survive the Flash-of-the-month cycle is context stored outside the model entirely.
That is the case for an external memory layer, stated plainly. MemX (memx.app) keeps notes, documents, preferences, and project context in a store the user owns, outside any single vendor's release cycle, so what an assistant knows on 3.7 Flash is still there on 3.8 and on whatever ships three weeks after that. It is private by architecture: per-user isolation, encryption at rest, on-device processing where possible, and no training on user data. Contrast that with the 3.8 Flash free tier, where 20 daily requests run under terms that allow training and human review, or with a cache that bills $0.50 per million tokens per hour to hold context. Memory you own costs nothing to keep warm and nothing to move.
Adopt now, stay on 3.7, or wait for January
There is no single right answer, because both halves of the story are true at once: the model is genuinely strong at agentic coding for its price, and it is genuinely more expensive to run per task than the model it follows. The decision splits by workload, and the introductory window makes timing part of the math. Whatever the choice, price the January cliff into any plan that runs past December.
- Adopt now: agentic coding and tool-use pipelines that can bank the introductory rate through December, especially via the Batch API at $0.375 in and $1.875 out
- Stay on 3.7 Flash: high-volume, efficiency-first workloads. Google recommends it in its own launch post, and Artificial Analysis's per-task figures, $0.40 versus $0.58, back it up
- Budget for January either way: any 2027 forecast should be modeled at $1.50 and $7.50 per million tokens, because the current rate is a promotion with a printed end date
- Keep sensitive work off the free tier: 20 requests a day under terms that permit training and human review is a trial experience, not an operating tier
- Do not build around a fixed model personality: on the current cadence, the next Flash is weeks away, and behavior tuned to this one will need re-testing
01Is Gemini 3.8 Flash free to use?
Barely. The consumer Gemini app shows it only to paid Google AI Pro and Ultra subscribers. The API free tier allows about 20 requests per day at launch, and Google's unpaid-services terms let it use those prompts and outputs to improve products, with possible human review.
02How much does Gemini 3.8 Flash cost?
Through December 31, 2026: $0.75 per million input tokens and $3.75 per million output tokens, with Batch API rates at half that. From January 1, 2027, prices double to $1.50 and $7.50, and batch, cached reads, and cache storage double on the same date.
03Why does Gemini 3.8 Flash pricing double in January 2027?
The launch rates are introductory. Google's pricing page lists $1.50 input and $7.50 output as the regular prices taking effect January 1, 2027, and applies the same schedule to Gemini 3.7 Flash and 3.6 Flash. All three sit on a promotion that expires December 31, 2026.
04How does Gemini 3.8 Flash compare to Claude Opus 5?
On three of Google's benchmark tables it matches or edges past Claude Opus 5: 10.0 against 6.7 on Harvey's Legal Agent, 73.7 against 74.0 on DeepSWE coding, 89.4 against 89.1 on Terminal-Bench 2.1, all at about a sixth of Opus 5's price. On the harder Terminal-Bench 4.0, Opus 5 still wins clearly, 51.8 against 19.1.
05Does Gemini 3.8 Flash have memory?
No. It is a stateless model with no personalization features of its own; memory lives in the Gemini app layer, not the model. Context caching can hold tokens temporarily, but it costs $0.50 per million tokens per hour to maintain and expires by design. Persistent context requires an external memory layer.
