[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"\u002Farticles\u002Fgemini-3-8-flash-vs-opus-5-benchmark?locale=en":3,"\u002Fcategories?locale=en":78},{"id":4,"slug":5,"type":6,"title":7,"excerpt":8,"coverImageUrl":9,"categoryId":10,"categorySlug":11,"categoryName":12,"categoryNameEn":13,"authorName":14,"status":15,"publishedAt":16,"seoTitle":17,"seoDescription":18,"ogImageUrl":9,"aiGenerated":19,"aiSourceNotes":20,"viewCount":21,"featured":22,"productsPosition":23,"locale":24,"translationGroupId":25,"unpublishedReason":14,"unpublishedAt":26,"createdAt":27,"updatedAt":28,"contentHtml":29,"plainSummary":18,"tags":30,"products":31,"related":32,"translations":74},"f4e4748d-4d0a-4cd9-92e2-b38b8c2c7b09","gemini-3-8-flash-vs-opus-5-benchmark","post","Gemini 3.8 Flash Drops: Official Benchmarks Tie or Beat Opus 5 on Most Tasks, but Two Big Gaps Are Impossible to Hide","Google launched Gemini 3.8 Flash on September 2, just three weeks after 3.7 Flash, priced at only one-sixth to one-seventh of Claude Opus 5. I broke down Google's official benchmark table into plain English: it ties or beats Opus 5 across most tasks, but the gap in Terminal-bench 4.0 and computer-use tests isn't nearly as small as the marketing pitch suggests.","\u002Fuploads\u002F1588b32d-8129-4f0f-a328-95e7613b44a9.jpg","df5d8a68-5464-4db3-b0ff-ce3843b48885","tech","科技","Tech","","published","2026-09-07T08:54:12.421Z","Gemini 3.8 Flash Benchmark Breakdown: Official Numbers vs. Opus 5 and GPT-5.6 Sol","A breakdown of Google's official Gemini 3.8 Flash benchmarks compared against Claude Opus 5 and GPT-5.6 Sol. At one-sixth the price, it holds its own—yet massive gaps remain in Terminal-bench 4.0 and OSWorld computer use. Includes Google AI Pro\u002FUltra subscriber access details and API cost calculations.",true,"翻譯自中文文章「Gemini 3.8 Flash 上線：官方數字曬出來，多數任務打平甚至贏 Opus 5，但這兩項差距藏不住」（原文 articleId: 39aebe51-103d-4029-b041-05aa0a7ab759）。推廣連結／產品卡尚未設定，英文版要用哪個聯盟計畫還沒決定，需要人工另外補上，不要照搬原文的台灣在地連結。分類與封面圖沿用原文，分類名稱目前還是中文顯示（categories 表還沒有英文名稱欄位，屬已知待補項目）。",0,false,"bottom","en","ccb5ac5b-0f0d-41c8-a0f8-0731c1026ebe",null,"2026-09-07T15:32:38.299Z","2026-09-07T16:54:12.421Z","\u003Cp>Google is moving fast this time. 3.6 Flash, 3.7 Flash, 3.8 Flash—this marks their third &quot;next-gen Flash&quot; announcement in three months, with this latest model dropping on September 2, just three weeks after 3.7 Flash. The Hacker News thread shot up to 775 points that day. What was unusual was that the debate wasn't about &quot;is it better than the last generation?&quot; but rather &quot;compared to flagship models a whole tier more expensive, where does it actually fall short?&quot; I broke down Google's official benchmark table into plain English, and the reality is a lot more honest than the marketing hype.\u003C\u002Fp>\n\u003Ch2>The Official Numbers: Tying or Even Beating Opus 5 on Most Tasks\u003C\u002Fh2>\n\u003Cp>First, let's talk pricing. 3.8 Flash runs $0.75 per million input tokens and $3.75 per million output tokens. That's introductory pricing good through December 31, 2026, after which it goes back up to $1.50 \u002F $7.50 starting January 1 next year. For comparison, Claude Opus 5 costs $5 \u002F $25, and GPT-5.6 Sol sits at $4 \u002F $20. Doing the math, 3.8 Flash currently costs roughly one-sixth to one-seventh of Opus 5.\u003C\u002Fp>\n\u003Cp>Given that price gap, most people would assume its performance would lag far behind Opus 5, but Google's official comparison table proves otherwise. On financial analysis with Vals Finance Agent v2, 3.8 Flash scores 61.4%, higher than Opus 5's 58.6%. On legal workflows with Harvey's Legal Agent Benchmark, 3.8 Flash hits 10.0% versus Opus 5's 6.7%. In long-horizon software engineering on DeepSWE v1.1, 3.8 Flash scores 73.7%, trailing Opus 5's 74.0% by just 0.3 percentage points—practically a dead heat. In terminal tasks on Terminal-bench 2.1, 3.8 Flash's 89.4% even edges past Opus 5's 89.1%. Across multimodal benchmarks—CharXiv Reasoning, long-video understanding on LVBench, and biological research on LABBench2—3.8 Flash is actually the top-scoring model on the entire board, beating every benchmarked rival, including Opus 5.\u003C\u002Fp>\n\u003Ch2>But in Two Areas, the Gap Is Too Big to Hide\u003C\u002Fh2>\n\u003Cp>Looking over the entire table, a clear pattern emerges: the tests it wins are all single-purpose tasks with well-defined scopes. But once you switch to tasks requiring long-horizon autonomous planning and operating an entire environment, 3.8 Flash reveals its true colors. The starkest example is Terminal-bench 4.0, which evaluates general agent capabilities under much more realistic scenarios: 3.8 Flash only manages 19.1%, while Opus 5 scores 51.8%—a massive 32.7 percentage-point loss and the single biggest deficit on the entire chart. Computer-use evaluation on OSWorld-2.0 tells the same story: 3.8 Flash hits 59.0% against Opus 5's 75.4%, lagging by 16.4 percentage points. On knowledge-work evaluation GDPVal-AA v2 (measured in Elo score), 3.8 Flash sits at 1545 while Opus 5 reaches 1824—another substantial spread.\u003C\u002Fp>\n\u003Cp>This discrepancy is actually consistent with Google's positioning in their official documentation. 3.8 Flash is pitched primarily for agent workloads with &quot;clear goals and relatively fixed steps,&quot; such as coding to spec or executing standardized financial and legal workflows. When it comes to open-ended tasks where the model has to figure out what to do next on its own and maintain long-term holistic planning, the flagship Opus 5 remains far more reliable. So claims of &quot;beating frontier models&quot; aren't technically lying—they're just cherry-picking the specific categories that happen to favor them.\u003C\u002Fp>\n\u003Ch2>Touted for \"Cost Savings,\" but Third-Party Usage Numbers Tell a Different Story\u003C\u002Fh2>\n\u003Cp>Google's selling point this time around is &quot;Flash speed and cost with flagship intelligence,&quot; and DeepMind's official page goes as far as claiming it is &quot;Best for token efficiency.&quot;\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"\u002Fuploads\u002F6d3e0351-6fa2-47e5-9c98-389021ab24b5.jpg\" alt=\"DeepMind's official page positions Gemini Flash as the best choice for token efficiency\">\u003C\u002Fp>\n\u003Cp>\u003Cem>DeepMind's official website positions Gemini Flash as &quot;the best choice for token efficiency across coding, knowledge work, and multimodal tasks.&quot; (Source: Google DeepMind official website)\u003C\u002Fem>\u003C\u002Fp>\n\u003Cp>However, real-world measurements from independent benchmarking firm Artificial Analysis show quite a contrast. At High reasoning intensity, 3.8 Flash's time to first token clocks in at 13.3 seconds, compared to a class median of 2.99 seconds—more than four times slower. Switching to Low reasoning drops latency to 0.7 seconds, but at the expense of giving up most of its reasoning horsepower. Output speed, on the other hand, is genuinely fast: in High mode, it churns out 304.6 tokens per second, blowing past the price-class median of 70.8 tokens per second. But spitting out tokens quickly doesn't mean it uses fewer of them. Across the same suite of benchmark tasks, 3.8 Flash consumed 120 million output tokens against a median of just 71 million—burning nearly 70% more tokens. Crunching the numbers, a cheaper per-token sticker price doesn't automatically translate to a cheaper overall bill for a completed run. Those are two very different metrics, and they need to be evaluated separately.\u003C\u002Fp>\n\u003Cp>There's also a regression that Google openly acknowledges. DeepMind's model card explicitly notes that 3.8 Flash slipped by 5.4 percentage points compared to 3.7 Flash in automated &quot;multilingual safety&quot; benchmarks. While human red-teaming still met launch thresholds, this means guardrails in non-English scenarios haven't improved—they've actually loosened slightly. Anyone building multilingual customer support or overseas-facing applications should keep this in mind.\u003C\u002Fp>\n\u003Ch2>Regular Subscribers Don't Need Add-Ons: Pro and Ultra Get Direct Access\u003C\u002Fh2>\n\u003Cp>Unlike what I covered in my recent piece on the \u003Ca href=\"\u002Farticles\u002Fclaude-fable-5-1-ai-subscription-comparison-2026\">Claude Fable 5.1 launch\u003C\u002Fa>—where even Pro subscribers couldn't use the model without paying for extra usage tiers—Google's official blog post is refreshingly straightforward. 3.8 Flash is already live for Google AI Pro and Ultra subscribers. No unlocking required, no extra fees; as long as you're subscribed, you can switch to this new model right away. Google AI Pro costs $19.99\u002Fmonth, while Ultra comes in two tiers: $99.99\u002Fmonth for 5x usage, and $199.99\u002Fmonth for 20x usage. All three tiers include 3.8 Flash directly within their base quota. For everyday users who just chat in the Gemini app and never touch an API, that's the most immediate difference.\u003C\u002Fp>\n\u003Cp>One detail to keep in mind, though: Google's own plan comparison page still promotes Gemini 3.1 Pro as its flagship offering. 3.8 Flash is an option you only see when switching models inside the app dropdown. Google's main marketing pages haven't really highlighted it, so it feels more like an under-the-radar bonus tucked into your subscription rather than the headliner of this rollout.\u003C\u002Fp>\n\u003Ch2>How Developers Can Use It—and What It Costs\u003C\u002Fh2>\n\u003Cp>Beyond the Gemini app, 3.8 Flash is simultaneously available across developer tooling like Google AI Studio, Android Studio, and Google Antigravity, while enterprise customers access it through Gemini Enterprise. As mentioned earlier, API pricing sits at the introductory $0.75 \u002F $3.75 rate through the end of this year. To put that into perspective for a single conversation turn—assuming 10,000 input tokens and 1,500 output tokens—you're looking at roughly $0.0075 plus $0.0056, coming out to about $0.013. That converts to less than NT$1 (New Taiwan Dollar), more than ten times cheaper than the extra-usage rates calculated in the Claude piece for Fable (which cost around $0.18 for the same token ratio). In terms of specs, 3.8 Flash supports text, image, audio, and video inputs, with a 1-million-token context window and a single-response output cap of 64k tokens—identical to the previous 3.7 Flash. DeepMind's model card makes no secret that this is continuous pre-training rather than a brand-new foundation model.\u003C\u002Fp>\n\u003Ch2>So, Should You Actually Care?\u003C\u002Fh2>\n\u003Cp>If you're just a Google AI Pro or Ultra subscriber using the Gemini app to ask questions and look up information, this update requires almost zero effort: just switch to the new model in the dropdown. In most situations, it'll be more accurate than 3.7 Flash, with the only real trade-off being occasionally slower response times. If you're building with the API—running agents, writing code, or wiring up structured workflows—3.8 Flash is undeniably compelling at this price point, and the official numbers back up the claim that it can punch above its weight and go toe-to-toe with flagships, provided the scope of the task is clear-cut. But the moment you venture into open-ended tasks demanding sustained autonomous planning and end-to-end computer control, the gaps in Terminal-bench 4.0 and OSWorld-2.0 speak for themselves. In those cases, you're still going to need Opus 5. Don't let flashy &quot;beats frontier models&quot; headlines sucker you into looking only at the numbers that happened to go Google's way.\u003C\u002Fp>\n",[],[],[33,47,61],{"id":34,"slug":35,"type":6,"title":36,"excerpt":37,"coverImageUrl":38,"contentMd":14,"categoryId":10,"categorySlug":11,"categoryName":12,"categoryNameEn":13,"authorName":14,"status":15,"publishedAt":39,"seoTitle":40,"seoDescription":41,"ogImageUrl":38,"aiGenerated":19,"aiSourceNotes":42,"viewCount":43,"featured":22,"productsPosition":23,"locale":24,"translationGroupId":44,"unpublishedReason":14,"unpublishedAt":26,"createdAt":45,"updatedAt":46},"b5c13d1d-606d-47bc-aaf9-cbaf8b12167c","apple-foldable-iphone-ultra-sept-10-event","Apple's Sept 10 Event Is Official: With a Rumored NT$60,000 Foldable iPhone Ultra, Should You Wait or Buy Current Stock Now?","Apple officially confirms its Sept 10 event at 1:00 AM. The first foldable iPhone is rumored to start around NT$63,000 for 256GB—roughly 1.6 times the current iPhone 17 Pro. Here is our breakdown of how to watch the livestream, whether the first-gen foldable is worth the gamble, and how to pick from in-stock models if you're on a budget.","\u002Fuploads\u002F1908d342-fada-4d13-8359-0ad938d43e95.jpg","2026-09-07T08:54:12.644Z","Apple Sept 10 Event Foldable iPhone Rumor Roundup: Wait for the New Model or Buy Current Stock?","Apple officially confirms its Sept 10 event at 1:00 AM. The first foldable iPhone is rumored around NT$63,000 for 256GB—1.6x the iPhone 17 Pro. We cover foldable spec rumors, supply risks, and how to choose between the iPhone 17, 16 Pro Max, and 15 Pro Max on a budget.","翻譯自中文文章「蘋果9\u002F10發表會確定登場，傳聞的摺疊iPhone Ultra六萬元預算，現在該等新機還是先衝現貨」（原文 articleId: 0f5c68b1-d575-4c76-aa32-64e5d9bba9cd）。推廣連結／產品卡尚未設定，英文版要用哪個聯盟計畫還沒決定，需要人工另外補上，不要照搬原文的台灣在地連結。分類與封面圖沿用原文，分類名稱目前還是中文顯示（categories 表還沒有英文名稱欄位，屬已知待補項目）。",1,"a84d999c-2357-4c29-8fc8-060134b901ad","2026-09-07T15:24:05.356Z","2026-09-07T16:54:12.644Z",{"id":48,"slug":49,"type":6,"title":50,"excerpt":51,"coverImageUrl":52,"contentMd":14,"categoryId":10,"categorySlug":11,"categoryName":12,"categoryNameEn":13,"authorName":14,"status":15,"publishedAt":53,"seoTitle":54,"seoDescription":55,"ogImageUrl":52,"aiGenerated":19,"aiSourceNotes":56,"viewCount":57,"featured":22,"productsPosition":23,"locale":24,"translationGroupId":58,"unpublishedReason":14,"unpublishedAt":26,"createdAt":59,"updatedAt":60},"3f366f84-f1ec-464d-a7f1-db1be8b0b6c5","chrome-zero-day-cve-2026-85046-edge-brave-update","Major Chrome Zero-Day Under Active Exploit: Edge and Brave Users Need to Update Too","Google confirmed that CVE-2026-85046 in Chrome's V8 engine was actively exploited in the wild before a patch was released. Rated at CVSS 8.8, this marks the sixth exploited Chrome zero-day this year. Because the flaw resides in the Chromium engine itself, browsers sharing the same foundation—like Edge and Brave—are all affected, and the Android version requires a separate update. Here is a walkthrough of how to check and update each browser.","\u002Fuploads\u002Fa768bdbc-efcc-4b81-976d-a2eb6f9a3e44.jpg","2026-09-07T08:54:12.590Z","How to Patch Chrome Zero-Day CVE-2026-85046: Edge, Brave, and Android Need Updates Too","On September 3, 2026, Google patched CVE-2026-85046 in Chrome's V8 engine, confirming active exploitation prior to the fix (CVSS 8.8). Here is a breakdown of patched versions, Edge and Brave status, Android vs. iPhone differences, and a 3-step guide to updating.","翻譯自中文文章「Chrome 出現正在被實際攻擊的重大零日漏洞! Edge、Brave 使用者也要一起更新」（原文 articleId: 3b33382e-bc5d-4367-b249-23609902f944）。推廣連結／產品卡尚未設定，英文版要用哪個聯盟計畫還沒決定，需要人工另外補上，不要照搬原文的台灣在地連結。分類與封面圖沿用原文，分類名稱目前還是中文顯示（categories 表還沒有英文名稱欄位，屬已知待補項目）。",4,"732cf87a-c84d-452a-aebd-6adcc76d49d8","2026-09-07T15:24:57.416Z","2026-09-07T16:54:12.590Z",{"id":62,"slug":63,"type":6,"title":64,"excerpt":65,"coverImageUrl":66,"contentMd":14,"categoryId":10,"categorySlug":11,"categoryName":12,"categoryNameEn":13,"authorName":14,"status":15,"publishedAt":67,"seoTitle":68,"seoDescription":69,"ogImageUrl":66,"aiGenerated":19,"aiSourceNotes":70,"viewCount":43,"featured":22,"productsPosition":23,"locale":24,"translationGroupId":71,"unpublishedReason":14,"unpublishedAt":26,"createdAt":72,"updatedAt":73},"794299c0-6021-46c5-b23f-d8cad171a645","line-account-deletion-facebook-phone-number","LINE Is Deleting Accounts by Year-End: Heads Up If You Only Linked Facebook Without a Phone Number!","The notices claiming \"LINE will delete accounts without a linked phone number by year-end\" are real, but they only target one specific group: accounts linked solely to Facebook with no registered phone number. Here is how to check if you're affected, how to set up your phone number, email, password, and Apple\u002FGoogle link all in one go, the major pitfall where linking a number deactivates another account, and the only path left if you're already locked out.","\u002Fuploads\u002F9d92f757-8b33-4121-bb37-2efb59510bb9.jpg","2026-09-07T08:54:12.534Z","LINE Year-End Account Deletions: How to Link Your Phone Number and Check Account Settings","Official LINE announcement: Accounts linked only to Facebook without a phone number must register a number before December 31, 2026, at 23:59, or they will be deleted. Learn how to verify your status, complete all four settings, avoid duplicate-binding pitfalls, and what to do if you're locked out.","翻譯自中文文章「LINE 年底要刪一批帳號，只連 FB 沒綁手機的人該注意了!」（原文 articleId: 5d261ee2-93b1-4d37-af75-2378fa5f4dbe）。推廣連結／產品卡尚未設定，英文版要用哪個聯盟計畫還沒決定，需要人工另外補上，不要照搬原文的台灣在地連結。分類與封面圖沿用原文，分類名稱目前還是中文顯示（categories 表還沒有英文名稱欄位，屬已知待補項目）。","af84222b-dcd4-4160-a2aa-a47c76d4412e","2026-09-07T15:28:31.921Z","2026-09-07T16:54:12.534Z",[75],{"locale":76,"slug":77},"zh-Hant","gemini-3-8-flash-benchmark-vs-opus-5-2026",[79,87,95,102,107],{"id":80,"slug":81,"name":82,"nameEn":83,"description":84,"descriptionEn":85,"sortOrder":43,"articleCount":86},"a7a26ffe-60e7-4548-b3d0-f271db563d51","travel","旅遊","Travel","訂房、機票、行程規劃相關的工具與心得","Booking, flights, and trip-planning tools and notes",2,{"id":88,"slug":89,"name":90,"nameEn":91,"description":92,"descriptionEn":93,"sortOrder":94,"articleCount":94},"5020f399-c954-48d2-bb8d-d330714ac6d6","finance","理財","Finance","記帳、投資、支付工具的使用經驗","Budgeting, investing, and payment tool experiences",3,{"id":96,"slug":97,"name":98,"nameEn":99,"description":100,"descriptionEn":101,"sortOrder":57,"articleCount":94},"418fe4fb-fbce-4f37-88e0-a917b7f2c87b","lifestyle","生活","Lifestyle","日常會用到的各種服務與小工具","Everyday services and small tools worth having",{"id":10,"slug":11,"name":12,"nameEn":13,"description":103,"descriptionEn":104,"sortOrder":105,"articleCount":106},"人類前進的動力","The engine behind human progress",5,7,{"id":108,"slug":109,"name":110,"nameEn":111,"description":112,"descriptionEn":113,"sortOrder":114,"articleCount":86},"84734e6c-fdfe-4544-b08a-0522c4df802f","game","遊戲","Games","身為玩家的最愛","A gamer's favorite things",6]