Gemini 3.8 Flash Drops: Official Benchmarks Tie or Beat Opus 5 on Most Tasks, but Two Big Gaps Are Impossible to Hide
Google launched Gemini 3.8 Flash on September 2, just three weeks after 3.7 Flash, priced at only one-sixth to one-seventh of Claude Opus 5. I broke down Google's official benchmark table into plain English: it ties or beats Opus 5 across most tasks, but the gap in Terminal-bench 4.0 and computer-use tests isn't nearly as small as the marketing pitch suggests.

Google is moving fast this time. 3.6 Flash, 3.7 Flash, 3.8 Flash—this marks their third "next-gen Flash" announcement in three months, with this latest model dropping on September 2, just three weeks after 3.7 Flash. The Hacker News thread shot up to 775 points that day. What was unusual was that the debate wasn't about "is it better than the last generation?" but rather "compared to flagship models a whole tier more expensive, where does it actually fall short?" I broke down Google's official benchmark table into plain English, and the reality is a lot more honest than the marketing hype.
The Official Numbers: Tying or Even Beating Opus 5 on Most Tasks
First, let's talk pricing. 3.8 Flash runs $0.75 per million input tokens and $3.75 per million output tokens. That's introductory pricing good through December 31, 2026, after which it goes back up to $1.50 / $7.50 starting January 1 next year. For comparison, Claude Opus 5 costs $5 / $25, and GPT-5.6 Sol sits at $4 / $20. Doing the math, 3.8 Flash currently costs roughly one-sixth to one-seventh of Opus 5.
Given that price gap, most people would assume its performance would lag far behind Opus 5, but Google's official comparison table proves otherwise. On financial analysis with Vals Finance Agent v2, 3.8 Flash scores 61.4%, higher than Opus 5's 58.6%. On legal workflows with Harvey's Legal Agent Benchmark, 3.8 Flash hits 10.0% versus Opus 5's 6.7%. In long-horizon software engineering on DeepSWE v1.1, 3.8 Flash scores 73.7%, trailing Opus 5's 74.0% by just 0.3 percentage points—practically a dead heat. In terminal tasks on Terminal-bench 2.1, 3.8 Flash's 89.4% even edges past Opus 5's 89.1%. Across multimodal benchmarks—CharXiv Reasoning, long-video understanding on LVBench, and biological research on LABBench2—3.8 Flash is actually the top-scoring model on the entire board, beating every benchmarked rival, including Opus 5.
But in Two Areas, the Gap Is Too Big to Hide
Looking over the entire table, a clear pattern emerges: the tests it wins are all single-purpose tasks with well-defined scopes. But once you switch to tasks requiring long-horizon autonomous planning and operating an entire environment, 3.8 Flash reveals its true colors. The starkest example is Terminal-bench 4.0, which evaluates general agent capabilities under much more realistic scenarios: 3.8 Flash only manages 19.1%, while Opus 5 scores 51.8%—a massive 32.7 percentage-point loss and the single biggest deficit on the entire chart. Computer-use evaluation on OSWorld-2.0 tells the same story: 3.8 Flash hits 59.0% against Opus 5's 75.4%, lagging by 16.4 percentage points. On knowledge-work evaluation GDPVal-AA v2 (measured in Elo score), 3.8 Flash sits at 1545 while Opus 5 reaches 1824—another substantial spread.
This discrepancy is actually consistent with Google's positioning in their official documentation. 3.8 Flash is pitched primarily for agent workloads with "clear goals and relatively fixed steps," such as coding to spec or executing standardized financial and legal workflows. When it comes to open-ended tasks where the model has to figure out what to do next on its own and maintain long-term holistic planning, the flagship Opus 5 remains far more reliable. So claims of "beating frontier models" aren't technically lying—they're just cherry-picking the specific categories that happen to favor them.
Touted for "Cost Savings," but Third-Party Usage Numbers Tell a Different Story
Google's selling point this time around is "Flash speed and cost with flagship intelligence," and DeepMind's official page goes as far as claiming it is "Best for token efficiency."

DeepMind's official website positions Gemini Flash as "the best choice for token efficiency across coding, knowledge work, and multimodal tasks." (Source: Google DeepMind official website)
However, real-world measurements from independent benchmarking firm Artificial Analysis show quite a contrast. At High reasoning intensity, 3.8 Flash's time to first token clocks in at 13.3 seconds, compared to a class median of 2.99 seconds—more than four times slower. Switching to Low reasoning drops latency to 0.7 seconds, but at the expense of giving up most of its reasoning horsepower. Output speed, on the other hand, is genuinely fast: in High mode, it churns out 304.6 tokens per second, blowing past the price-class median of 70.8 tokens per second. But spitting out tokens quickly doesn't mean it uses fewer of them. Across the same suite of benchmark tasks, 3.8 Flash consumed 120 million output tokens against a median of just 71 million—burning nearly 70% more tokens. Crunching the numbers, a cheaper per-token sticker price doesn't automatically translate to a cheaper overall bill for a completed run. Those are two very different metrics, and they need to be evaluated separately.
There's also a regression that Google openly acknowledges. DeepMind's model card explicitly notes that 3.8 Flash slipped by 5.4 percentage points compared to 3.7 Flash in automated "multilingual safety" benchmarks. While human red-teaming still met launch thresholds, this means guardrails in non-English scenarios haven't improved—they've actually loosened slightly. Anyone building multilingual customer support or overseas-facing applications should keep this in mind.
Regular Subscribers Don't Need Add-Ons: Pro and Ultra Get Direct Access
Unlike what I covered in my recent piece on the Claude Fable 5.1 launch—where even Pro subscribers couldn't use the model without paying for extra usage tiers—Google's official blog post is refreshingly straightforward. 3.8 Flash is already live for Google AI Pro and Ultra subscribers. No unlocking required, no extra fees; as long as you're subscribed, you can switch to this new model right away. Google AI Pro costs $19.99/month, while Ultra comes in two tiers: $99.99/month for 5x usage, and $199.99/month for 20x usage. All three tiers include 3.8 Flash directly within their base quota. For everyday users who just chat in the Gemini app and never touch an API, that's the most immediate difference.
One detail to keep in mind, though: Google's own plan comparison page still promotes Gemini 3.1 Pro as its flagship offering. 3.8 Flash is an option you only see when switching models inside the app dropdown. Google's main marketing pages haven't really highlighted it, so it feels more like an under-the-radar bonus tucked into your subscription rather than the headliner of this rollout.
How Developers Can Use It—and What It Costs
Beyond the Gemini app, 3.8 Flash is simultaneously available across developer tooling like Google AI Studio, Android Studio, and Google Antigravity, while enterprise customers access it through Gemini Enterprise. As mentioned earlier, API pricing sits at the introductory $0.75 / $3.75 rate through the end of this year. To put that into perspective for a single conversation turn—assuming 10,000 input tokens and 1,500 output tokens—you're looking at roughly $0.0075 plus $0.0056, coming out to about $0.013. That converts to less than NT$1 (New Taiwan Dollar), more than ten times cheaper than the extra-usage rates calculated in the Claude piece for Fable (which cost around $0.18 for the same token ratio). In terms of specs, 3.8 Flash supports text, image, audio, and video inputs, with a 1-million-token context window and a single-response output cap of 64k tokens—identical to the previous 3.7 Flash. DeepMind's model card makes no secret that this is continuous pre-training rather than a brand-new foundation model.
So, Should You Actually Care?
If you're just a Google AI Pro or Ultra subscriber using the Gemini app to ask questions and look up information, this update requires almost zero effort: just switch to the new model in the dropdown. In most situations, it'll be more accurate than 3.7 Flash, with the only real trade-off being occasionally slower response times. If you're building with the API—running agents, writing code, or wiring up structured workflows—3.8 Flash is undeniably compelling at this price point, and the official numbers back up the claim that it can punch above its weight and go toe-to-toe with flagships, provided the scope of the task is clear-cut. But the moment you venture into open-ended tasks demanding sustained autonomous planning and end-to-end computer control, the gaps in Terminal-bench 4.0 and OSWorld-2.0 speak for themselves. In those cases, you're still going to need Opus 5. Don't let flashy "beats frontier models" headlines sucker you into looking only at the numbers that happened to go Google's way.
Sponsored
Related

Apple's Sept 10 Event Is Official: With a Rumored NT$60,000 Foldable iPhone Ultra, Should You Wait or Buy Current Stock Now?
Apple officially confirms its Sept 10 event at 1:00 AM. The first foldable iPhone is rumored to start around NT$63,000 for 256GB—roughly 1.6 times the current iPhone 17 Pro. Here is our breakdown of how to watch the livestream, whether the first-gen foldable is worth the gamble, and how to pick from in-stock models if you're on a budget.

Major Chrome Zero-Day Under Active Exploit: Edge and Brave Users Need to Update Too
Google confirmed that CVE-2026-85046 in Chrome's V8 engine was actively exploited in the wild before a patch was released. Rated at CVSS 8.8, this marks the sixth exploited Chrome zero-day this year. Because the flaw resides in the Chromium engine itself, browsers sharing the same foundation—like Edge and Brave—are all affected, and the Android version requires a separate update. Here is a walkthrough of how to check and update each browser.

LINE Is Deleting Accounts by Year-End: Heads Up If You Only Linked Facebook Without a Phone Number!
The notices claiming "LINE will delete accounts without a linked phone number by year-end" are real, but they only target one specific group: accounts linked solely to Facebook with no registered phone number. Here is how to check if you're affected, how to set up your phone number, email, password, and Apple/Google link all in one go, the major pitfall where linking a number deactivates another account, and the only path left if you're already locked out.