Model review · Official benchmarks · Same-prompt API tests
Gemini 3.8 Flash is Google’s new general-availability Flash model for long-horizon software engineering, agents, and complex knowledge work. The launch claims are strong. Our first controlled tests are much less tidy.
This Gemini 3.8 Flash review separates official facts from hands-on evidence. Google’s pages establish the release date, model status, context window, pricing, tools, and benchmark claims. Seven same-prompt API tasks show what the model actually returned beside Gemini 3.7 Flash.
If you want to repeat the comparison without managing separate model tabs, GlobalGPT provides one workspace for testing multiple AI models. For context on the previous generation, our Gemini 3.6 闪评 shows how quickly the Flash line has been changing.
Gemini 3.8 Flash Review: The Short Answer
Gemini 3.8 Flash did not beat Gemini 3.7 Flash in this first seven-task review. Across the six usable pairs, 3.8 won the frontend planning task, tied on a Python bug fix, and lost four tasks on status precision, data reasoning, bounded agent planning, and SEO editing. The remaining research task was inconclusive.
The gap was not about raw knowledge. It came from small defects in finished work: one wrong denominator claim, looser status labels, less economical prose, and a less concrete tool plan. Those are exactly the details that affect whether an answer can be used immediately.
评论快照
| 度量 | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| 任务胜利 | 1 | 4 |
| Ties | 1 | 1 |
| Completed-task elapsed time | 49.266s | 42.909s |
| 已完成任务的输出令牌 | 9,199 | 7,489 |
| Estimated completed-task cost | $0.036121 | $0.028569 |
These results describe six small text tasks on one API route and one configuration. They are not a universal ranking of either model.
What Is Gemini 3.8 Flash?
Google describes Gemini 3.8 Flash as its most intelligent Flash workhorse for coding and agents, with particular emphasis on long-horizon software engineering and knowledge work. It is based on Gemini 3.7 Flash rather than a completely separate architecture.
The positioning matters because this is not merely a cheaper general chatbot. Google is pitching 3.8 Flash as a production agent model that can reason across documents, tools, code, and longer task sequences while preserving Flash-level latency and scale.

Gemini 3.8 Flash Release Date
"(《世界人权宣言》) 官方人物卡 was published on September 2, 2026. Google’s product page lists the model as generally available rather than preview-only.
What Changed From Gemini 3.7 Flash?
The official emphasis shifts toward longer-running agent tasks and finished artifacts. Google highlights software engineering, rigorous reasoning, tool execution, multimodal understanding, and complex enterprise work. In our small test pack, that positioning translated into the strongest frontend plan, but not more reliable finished answers across the board.
Gemini 3.8 Flash Specs and Availability
| 规格 | Gemini 3.8 Flash | 证据 |
|---|---|---|
| API 模型 ID | gemini-3.8-flash | Official pricing page and returned API responses |
| 现状 | General availability | Google DeepMind product page |
| 输入上下文 | 最多 100 万个代币 | 官方模特卡 |
| 最大文本输出量 | 64K 个代币 | 官方模特卡 |
| Inputs | Text, images, audio, video, and PDFs | 官方产品页面 |
| 输出 | 文本 | 官方模特卡 |
| 工具支持 | Function calling, Search as a tool, computer use | 官方产品页面 |
| Official access surfaces | Gemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google Antigravity | 官方模特卡 |
| 知识分数线 | March 2026 for some domains; January 2025 for others | 官方模特卡 |
A 1M-token window is capacity, not proof of perfect long-document recall. The test pack below includes a short status-sensitive extraction task, but it does not establish million-token performance.
Gemini 3.8 Flash Pricing and API Cost
"(《世界人权宣言》) Gemini API 官方定价页面 lists a launch rate through December 31, 2026 of 每100万个输入代币为$0.75 和 每100万个输出代币为$3.75, including thinking tokens. On January 1, 2027, those rates rise to $1.50 and $7.50.
Published price change
Paid-tier price per 1M tokens. The January 2027 rates are double the introductory 2026 rates.
Gemini 3.8 Flash and Gemini 3.7 Flash had the same published introductory rate in our test period. Even so, 3.8 cost more in every completed pair because its recorded token usage was higher. Across T02–T07, the estimates were $0.036121 for 3.8 and $0.028569 for 3.7.
For a different point on the speed-and-cost curve, see our Gemini 3.5 Flash Lite review. A Lite model serves a different workload, so its lower price should not be treated as an automatic quality win.
Gemini 3.8 Flash Official Benchmarks
Google’s performance page emphasizes four areas: DeepSWE v1.1, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, and HLE-Verified. Its charts show Gemini 3.8 Flash ahead of Gemini 3.7 Flash on all three evaluations with exact published percentages.
| 官方评估 | Gemini 3.8 Flash | Gemini 3.7 Flash | 证据边界 |
|---|---|---|---|
| DeepSWE v1.1 | Above 70% in Google’s chart | Lower than 3.8 in the chart | Google says 3.8 leads in the efficient area; the public graphic does not expose a precise numeric label. |
| Vals Finance Agent v2 | 61.4% | 59.0% | Official Google chart; the benchmark’s full methodology still matters. |
| Harvey’s Legal Agent Benchmark | 10.0% | 8.8% | This benchmark result is not a guarantee of legal accuracy in deployment. |
| HLE-Verified | 54.9% | 53.6% | It does not directly measure SEO editing, UI planning, or response economy. |

How We Tested Gemini 3.8 Flash
Both models received the same task prompt, temperature of 0, a 3,200-token output ceiling, and the same API route. T02 was rerun once for both models after the first 3.8 request failed technically; only the matched rerun is scored.
| 记录字段 | How it was used |
|---|---|
| Returned model ID | Confirmed that the requested model answered. |
| Real output | Compared correctness, completeness, status fidelity, and editing required. |
| 已过去的时间(秒) | Measured local end-to-end response time, not provider-side latency. |
| 输入和输出令牌 | Supported task-level cost estimates at the official introductory rate. |
| Finish reason | Separated usable answers from incomplete or budget-limited output. |
Gemini 3.8 Flash Real-World Test Results
Each card keeps the task, the actual returned work, the measurements, and the editorial verdict together. The excerpts are shortened for readability but preserve the deciding details.
Test 1: Research Synthesis
任务: Produce a 180-word buying recommendation from a locked fact set, with three reasons, two unknowns, and a test plan.
Gemini 3.8 Flash
18.211s · 361 input / 3,196 output tokens · est. $0.012256
“We recommend selecting Model A… 3. Cost Parity: The introductory API price is the same for Model A and Model B”
Gemini 3.7 Flash
13.490s · 108 input / 3,196 output tokens · est. $0.012066
“We directly recommend piloting and purchasing Model A… The introductory API price remains the exact same for both models through December 3”
Both answers stop mid-sentence before the required unknowns and test plan. The 3.8 response reports 停止, while 3.7 reports 长度; neither finish reason changes the usability judgment.
Test 2: Python Bug Fix
任务: Return the three highest-value paid records, preserve input order on ties, handle empty input, explain the root cause, and add two tests.
Gemini 3.8 Flash
7.969s · 344 input / 1,517 output tokens · est. $0.005947
paid.sort(key=lambda r: r["amount"], reverse=True)
It explained Python’s stable sort and tested ranking, ties, filtering, and empty input.
Gemini 3.7 Flash
7.818s · 91 input / 1,359 output tokens · est. $0.005164
paid.sort(key=lambda r: r["amount"], reverse=True)
It returned the same minimal fix and covered the same acceptance criteria.
The code is functionally equivalent. The 3.7 answer is slightly shorter and cheaper, but that is not enough to separate output quality.
Test 3: Textual Frontend Implementation Plan
任务: Plan five prioritized fixes, responsive behavior, and QA for a described checkout. No screenshot was supplied.
Gemini 3.8 Flash
10.128s · 340 input / 1,509 output tokens · est. $0.005914
It specified a 58/42 desktop grid, 48–52px CTA, nearby trust note, contextual discount alert, mobile order-summary accordion, focus states, and overflow checks.
Gemini 3.7 Flash
8.149s · 87 input / 1,416 output tokens · est. $0.005375
It proposed a similar grid and mobile stack, but introduced “256-bit Encrypted Checkout,” a detail absent from the task.
Both plans are implementable. The 3.8 plan has the stronger QA checklist and avoids inventing a specific encryption claim.
Test 4: Status-Aware Extraction
任务: Summarize Orion and preserve confirmed, planned, approved, pending, unknown, and unverified states in a fact table.
Gemini 3.8 Flash
6.292s · 371 input / 1,122 output tokens · est. $0.004486
It labeled the second pilot “Confirmed / Pending” and wrote that the first pilot “successfully handled” 82% of cases.
Gemini 3.7 Flash
7.126s · 118 input / 1,152 output tokens · est. $0.004409
It kept the rollout “Planned (Unapproved),” legal review “Planned (Pending),” and production readiness “Unknown.”
Both retain the core facts, but 3.7 separates approval, planning, pending review, and unknown status more precisely. Both add “successfully” to the neutral 82% pilot statement.
Test 5: Data Reasoning
任务: Calculate ROAS and paid conversion rate, handle a zero-spend denominator correctly, rank channels, and recommend a conditional $300 split.
Gemini 3.8 Flash
11.074s · 377 input / 2,355 output tokens · est. $0.009114
It correctly ranked Email 7.0×, Search 3.0×, Social 2.6×, and excluded Partner ROAS. It then said zero spend creates a zero denominator for both ROAS and CPA.
Gemini 3.7 Flash
9.106s · 124 input / 1,969 output tokens · est. $0.007477
It returned the same correct rankings and treated Partner ROAS as not measurable, then proposed $200 Email / $100 Search when email can scale.
The 3.8 answer makes a real mathematical error: CPA is spend divided by conversions, so $0 divided by 4 is $0, not a zero-denominator case.
Test 6: Bounded Agentic Planning
任务: Use four named tools to collect one official API price, stop on unofficial or ambiguous evidence, and define completion in under 220 words.
Gemini 3.8 Flash
5.603s · 354 input / 948 output tokens · est. $0.003821
It mapped search, fetch, parse, and persist steps, then added authenticity, availability, and price-clarity gates.
Gemini 3.7 Flash
5.127s · 101 input / 894 output tokens · est. $0.003428
It used explicit tool-call syntax, named model_pricing.csv, required distinct input and output rates, and defined one verified row as completion.
Both plans respect the stop conditions. The 3.7 plan is more concrete about the available calls and the saved artifact without pretending the tools ran.
Test 7: SEO Writing and Editing
任务: Return HTML only, use “Gemini 3.8 Flash review” exactly once, preserve four locked facts, and avoid unsupported results, prices, dates, or access claims.
Gemini 3.8 Flash
8.200s · 379 input / 1,748 output tokens · est. $0.006839
<p>Our Gemini 3.8 Flash review evaluates the Google model by comparing official benchmark context with same-prompt API tests, with hands-on results pending as the evaluation records:</p>Gemini 3.7 Flash
5.583s · 126 input / 699 output tokens · est. $0.002716
<p>This Gemini 3.8 Flash review examines the new Google model by comparing official benchmark context with same-prompt API tests, with hands-on results currently pending.</p>Both satisfy the structure and locked facts. The 3.7 sentence is cleaner and reads like publishable SEO copy; 3.8’s “as the evaluation records” construction is awkward.
Gemini 3.8 Flash Test Scorecard
| 任务 | 优胜者 | Deciding difference |
|---|---|---|
| 研究综述 | 尚无定论 | Both incomplete |
| Python bug fix | 领带 | Equivalent code and coverage |
| Frontend plan | Gemini 3.8 Flash | Better QA and no invented encryption claim |
| Status extraction | Gemini 3.7 Flash | More precise state labels |
| 数据推理 | Gemini 3.7 Flash | No CPA denominator error |
| 代理规划 | Gemini 3.7 Flash | More concrete calls and artifact |
| SEO编辑 | Gemini 3.7 Flash | Cleaner publishable sentence |
The result is a useful counterweight to the launch story. Gemini 3.8 Flash may be stronger on the official agent and long-horizon evaluations Google chose to highlight, yet 3.7 produced better finished work in four of our small controlled tasks.
Gemini 3.8 Flash Limitations We Observed
The biggest issue was not a dramatic failure. It was extra correction work: a denominator mistake, vague status labeling, an awkward SEO sentence, and more tokens used on every completed task. T01 also ended before completing a tightly specified 180-word deliverable.
This test pack used text only. It does not evaluate image, audio, video, PDFs, computer use, true long-context retrieval, or a multi-hour agent run. Google’s model card also notes possible hallucinations, occasional slowness or timeouts, and higher token use at stronger effort levels.
Gemini 3.8 Flash Review: Final Verdict
Gemini 3.8 Flash is a meaningful official upgrade for agentic and long-horizon workloads, but it is not an automatic upgrade for every everyday task. It won one of six usable pairs in this review, while Gemini 3.7 Flash won four and used less time, fewer output tokens, and lower estimated cost.
The practical lesson is to test the work you actually ship. If your workload is long-running software engineering or document-heavy agent execution, Google’s official benchmark direction makes 3.8 worth a serious trial. If your priority is concise editing, small data analyses, or status-sensitive summaries, our first results say 3.7 remains highly competitive.
The same pattern appeared in our earlier Gemini 3.6 Flash 与 Gemini 3.1 Pro 对比: generation labels do not replace side-by-side output inspection. Run a recurring prompt in GlobalGPT, inspect the finished answer, and choose from evidence rather than the version number.
Gemini 3.8 Flash FAQ
When was Gemini 3.8 Flash released?
Google DeepMind published the Gemini 3.8 Flash model card on September 2, 2026.
Is Gemini 3.8 Flash generally available?
Yes. Google’s official product page lists Gemini 3.8 Flash as generally available.
What is the Gemini 3.8 Flash API model ID?
The official model ID is gemini-3.8-flash, which also matched the returned model ID in our API tests.
How much does Gemini 3.8 Flash cost?
Through December 31, 2026, Google lists $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, the rates rise to $1.50 and $7.50.
Is Gemini 3.8 Flash free?
Google’s pricing page lists a free tier with limited model access and free input and output tokens. Limits and model availability can differ by account and route.
What is the Gemini 3.8 Flash context window?
Gemini 3.8 Flash supports up to 1M input tokens and up to 64K tokens of text output, according to the official model card.
Does Gemini 3.8 Flash support images, audio, video, and PDFs?
Yes as inputs. Google’s product page lists text, image, video, audio, and PDF inputs; the model card lists text output.
Does Gemini 3.8 Flash support tools?
Google lists function calling, Search as a tool, and computer use. Tool availability can depend on the product and API route.
Is Gemini 3.8 Flash better than Gemini 3.7 Flash?
Not in every task. Gemini 3.8 Flash won one of six usable task pairs in this review, Gemini 3.7 Flash won four, and one pair tied. Google’s official evaluations emphasize stronger long-horizon agent and software-engineering performance.
Is Gemini 3.8 Flash good for coding?
Both models produced the correct minimal Python fix in our coding task. Google’s official positioning and DeepSWE claim make 3.8 especially relevant to longer software-engineering workflows, which our small bug fix did not measure.
Why did Gemini 3.8 Flash cost more in the tests if the published rate was the same?
Its recorded token usage was higher in every completed pair. A shared per-token rate can still produce a higher total cost when one response consumes more tokens.
Where can I access Gemini 3.8 Flash?
Google lists the Gemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google Antigravity. Availability can still vary by account, product, and region.
Are the benchmark results independently verified here?
No. The benchmark section attributes Google’s official claims directly. Artificial Analysis was unreachable during research, so its leaderboard data is not included as verified evidence in this draft.



