Model review · Official benchmarks · Same-prompt API tests
Gemini 3.8 Flash is Google’s new general-availability Flash model for long-horizon software engineering, agents, and complex knowledge work. The launch claims are strong. Our first controlled tests are much less tidy.
This Gemini 3.8 Flash review separates official facts from hands-on evidence. Google’s pages establish the release date, model status, context window, pricing, tools, and benchmark claims. Seven same-prompt API tasks show what the model actually returned beside Gemini 3.7 Flash.
If you want to repeat the comparison without managing separate model tabs, GlobalGPT provides one workspace for testing multiple AI models. For context on the previous generation, our Gemini 3.6 : aperçu rapide shows how quickly the Flash line has been changing.
Gemini 3.8 Flash Review: The Short Answer
Gemini 3.8 Flash did not beat Gemini 3.7 Flash in this first seven-task review. Across the six usable pairs, 3.8 won the frontend planning task, tied on a Python bug fix, and lost four tasks on status precision, data reasoning, bounded agent planning, and SEO editing. The remaining research task was inconclusive.
The gap was not about raw knowledge. It came from small defects in finished work: one wrong denominator claim, looser status labels, less economical prose, and a less concrete tool plan. Those are exactly the details that affect whether an answer can be used immediately.
Aperçu de l'avis
| Mesure | Gemini 3.8 Flash | Gemini 3,7 Flash |
|---|---|---|
| Victoires par tâche | 1 | 4 |
| Ties | 1 | 1 |
| Completed-task elapsed time | 49.266s | 42.909s |
| Jetons de sortie correspondant aux tâches terminées | 9,199 | 7,489 |
| Estimated completed-task cost | $0.036121 | $0.028569 |
These results describe six small text tasks on one API route and one configuration. They are not a universal ranking of either model.
What Is Gemini 3.8 Flash?
Google describes Gemini 3.8 Flash as its most intelligent Flash workhorse for coding and agents, with particular emphasis on long-horizon software engineering and knowledge work. It is based on Gemini 3.7 Flash rather than a completely separate architecture.
The positioning matters because this is not merely a cheaper general chatbot. Google is pitching 3.8 Flash as a production agent model that can reason across documents, tools, code, and longer task sequences while preserving Flash-level latency and scale.

Gemini 3.8 Flash Release Date
Le fiche officielle du modèle was published on September 2, 2026. Google’s product page lists the model as generally available rather than preview-only.
What Changed From Gemini 3.7 Flash?
The official emphasis shifts toward longer-running agent tasks and finished artifacts. Google highlights software engineering, rigorous reasoning, tool execution, multimodal understanding, and complex enterprise work. In our small test pack, that positioning translated into the strongest frontend plan, but not more reliable finished answers across the board.
Gemini 3.8 Flash Specs and Availability
| Spécification | Gemini 3.8 Flash | Preuves |
|---|---|---|
| ID du modèle API | gemini-3.8-flash | Official pricing page and returned API responses |
| Statut | General availability | Google DeepMind product page |
| Contexte de saisie | Jusqu'à 1 million de jetons | Fiche officielle du modèle |
| Longueur maximale du texte affiché | 64 000 jetons | Fiche officielle du modèle |
| Inputs | Text, images, audio, video, and PDFs | Page officielle du produit |
| Sortie | Texte | Fiche officielle du modèle |
| Assistance technique | Function calling, Search as a tool, computer use | Page officielle du produit |
| Official access surfaces | Gemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google Antigravity | Fiche officielle du modèle |
| Critères d'évaluation des connaissances | March 2026 for some domains; January 2025 for others | Fiche officielle du modèle |
A 1M-token window is capacity, not proof of perfect long-document recall. The test pack below includes a short status-sensitive extraction task, but it does not establish million-token performance.
Gemini 3.8 Flash Pricing and API Cost
Le Page officielle des tarifs de l'API Gemini lists a launch rate through December 31, 2026 of 1 TP41T0,75 par million de jetons d'entrée et 1 TP41T3,75 par million de jetons émis, including thinking tokens. On January 1, 2027, those rates rise to $1.50 and $7.50.
Published price change
Paid-tier price per 1M tokens. The January 2027 rates are double the introductory 2026 rates.
Gemini 3.8 Flash and Gemini 3.7 Flash had the same published introductory rate in our test period. Even so, 3.8 cost more in every completed pair because its recorded token usage was higher. Across T02–T07, the estimates were $0.036121 for 3.8 and $0.028569 for 3.7.
For a different point on the speed-and-cost curve, see our Gemini 3.5 Flash Lite review. A Lite model serves a different workload, so its lower price should not be treated as an automatic quality win.
Gemini 3.8 Flash Official Benchmarks
Google’s performance page emphasizes four areas: DeepSWE v1.1, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, and HLE-Verified. Its charts show Gemini 3.8 Flash ahead of Gemini 3.7 Flash on all three evaluations with exact published percentages.
| Évaluation officielle | Gemini 3.8 Flash | Gemini 3,7 Flash | Limites de la preuve |
|---|---|---|---|
| DeepSWE v1.1 | Above 70% in Google’s chart | Lower than 3.8 in the chart | Google says 3.8 leads in the efficient area; the public graphic does not expose a precise numeric label. |
| Vals Finance Agent v2 | 61.4% | 59.0% | Official Google chart; the benchmark’s full methodology still matters. |
| Harvey’s Legal Agent Benchmark | 10.0% | 8.8% | This benchmark result is not a guarantee of legal accuracy in deployment. |
| HLE-Verified | 54.9% | 53.6% | It does not directly measure SEO editing, UI planning, or response economy. |

How We Tested Gemini 3.8 Flash
Both models received the same task prompt, temperature of 0, a 3,200-token output ceiling, and the same API route. T02 was rerun once for both models after the first 3.8 request failed technically; only the matched rerun is scored.
| Champ enregistré | How it was used |
|---|---|
| Returned model ID | Confirmed that the requested model answered. |
| Real output | Compared correctness, completeness, status fidelity, and editing required. |
| Nombre de secondes écoulées | Measured local end-to-end response time, not provider-side latency. |
| Jetons d'entrée et de sortie | Supported task-level cost estimates at the official introductory rate. |
| Finish reason | Separated usable answers from incomplete or budget-limited output. |
Gemini 3.8 Flash Real-World Test Results
Each card keeps the task, the actual returned work, the measurements, and the editorial verdict together. The excerpts are shortened for readability but preserve the deciding details.
Test 1: Research Synthesis
Tâche : Produce a 180-word buying recommendation from a locked fact set, with three reasons, two unknowns, and a test plan.
Gemini 3.8 Flash
18.211s · 361 input / 3,196 output tokens · est. $0.012256
“We recommend selecting Model A… 3. Cost Parity: The introductory API price is the same for Model A and Model B”
Gemini 3,7 Flash
13.490s · 108 input / 3,196 output tokens · est. $0.012066
“We directly recommend piloting and purchasing Model A… The introductory API price remains the exact same for both models through December 3”
Both answers stop mid-sentence before the required unknowns and test plan. The 3.8 response reports arrêt, while 3.7 reports longueur; neither finish reason changes the usability judgment.
Test 2: Python Bug Fix
Tâche : Return the three highest-value paid records, preserve input order on ties, handle empty input, explain the root cause, and add two tests.
Gemini 3.8 Flash
7.969s · 344 input / 1,517 output tokens · est. $0.005947
paid.sort(key=lambda r: r["amount"], reverse=True)
It explained Python’s stable sort and tested ranking, ties, filtering, and empty input.
Gemini 3,7 Flash
7.818s · 91 input / 1,359 output tokens · est. $0.005164
paid.sort(key=lambda r: r["amount"], reverse=True)
It returned the same minimal fix and covered the same acceptance criteria.
The code is functionally equivalent. The 3.7 answer is slightly shorter and cheaper, but that is not enough to separate output quality.
Test 3: Textual Frontend Implementation Plan
Tâche : Plan five prioritized fixes, responsive behavior, and QA for a described checkout. No screenshot was supplied.
Gemini 3.8 Flash
10.128s · 340 input / 1,509 output tokens · est. $0.005914
It specified a 58/42 desktop grid, 48–52px CTA, nearby trust note, contextual discount alert, mobile order-summary accordion, focus states, and overflow checks.
Gemini 3,7 Flash
8.149s · 87 input / 1,416 output tokens · est. $0.005375
It proposed a similar grid and mobile stack, but introduced “256-bit Encrypted Checkout,” a detail absent from the task.
Both plans are implementable. The 3.8 plan has the stronger QA checklist and avoids inventing a specific encryption claim.
Test 4: Status-Aware Extraction
Tâche : Summarize Orion and preserve confirmed, planned, approved, pending, unknown, and unverified states in a fact table.
Gemini 3.8 Flash
6.292s · 371 input / 1,122 output tokens · est. $0.004486
It labeled the second pilot “Confirmed / Pending” and wrote that the first pilot “successfully handled” 82% of cases.
Gemini 3,7 Flash
7.126s · 118 input / 1,152 output tokens · est. $0.004409
It kept the rollout “Planned (Unapproved),” legal review “Planned (Pending),” and production readiness “Unknown.”
Both retain the core facts, but 3.7 separates approval, planning, pending review, and unknown status more precisely. Both add “successfully” to the neutral 82% pilot statement.
Test 5: Data Reasoning
Tâche : Calculate ROAS and paid conversion rate, handle a zero-spend denominator correctly, rank channels, and recommend a conditional $300 split.
Gemini 3.8 Flash
11.074s · 377 input / 2,355 output tokens · est. $0.009114
It correctly ranked Email 7.0×, Search 3.0×, Social 2.6×, and excluded Partner ROAS. It then said zero spend creates a zero denominator for both ROAS and CPA.
Gemini 3,7 Flash
9.106s · 124 input / 1,969 output tokens · est. $0.007477
It returned the same correct rankings and treated Partner ROAS as not measurable, then proposed $200 Email / $100 Search when email can scale.
The 3.8 answer makes a real mathematical error: CPA is spend divided by conversions, so $0 divided by 4 is $0, not a zero-denominator case.
Test 6: Bounded Agentic Planning
Tâche : Use four named tools to collect one official API price, stop on unofficial or ambiguous evidence, and define completion in under 220 words.
Gemini 3.8 Flash
5.603s · 354 input / 948 output tokens · est. $0.003821
It mapped search, fetch, parse, and persist steps, then added authenticity, availability, and price-clarity gates.
Gemini 3,7 Flash
5.127s · 101 input / 894 output tokens · est. $0.003428
It used explicit tool-call syntax, named model_pricing.csv, required distinct input and output rates, and defined one verified row as completion.
Both plans respect the stop conditions. The 3.7 plan is more concrete about the available calls and the saved artifact without pretending the tools ran.
Test 7: SEO Writing and Editing
Tâche : Return HTML only, use “Gemini 3.8 Flash review” exactly once, preserve four locked facts, and avoid unsupported results, prices, dates, or access claims.
Gemini 3.8 Flash
8.200s · 379 input / 1,748 output tokens · est. $0.006839
<p>Our Gemini 3.8 Flash review evaluates the Google model by comparing official benchmark context with same-prompt API tests, with hands-on results pending as the evaluation records:</p>Gemini 3,7 Flash
5.583s · 126 input / 699 output tokens · est. $0.002716
<p>This Gemini 3.8 Flash review examines the new Google model by comparing official benchmark context with same-prompt API tests, with hands-on results currently pending.</p>Both satisfy the structure and locked facts. The 3.7 sentence is cleaner and reads like publishable SEO copy; 3.8’s “as the evaluation records” construction is awkward.
Gemini 3.8 Flash Test Scorecard
| Tâche | Gagnant | Deciding difference |
|---|---|---|
| Synthèse de recherche | Non concluant | Both incomplete |
| Python bug fix | Cravate | Equivalent code and coverage |
| Frontend plan | Gemini 3.8 Flash | Better QA and no invented encryption claim |
| Status extraction | Gemini 3,7 Flash | More precise state labels |
| Raisonnement à partir des données | Gemini 3,7 Flash | No CPA denominator error |
| Planification agentique | Gemini 3,7 Flash | More concrete calls and artifact |
| Rédaction SEO | Gemini 3,7 Flash | Cleaner publishable sentence |
The result is a useful counterweight to the launch story. Gemini 3.8 Flash may be stronger on the official agent and long-horizon evaluations Google chose to highlight, yet 3.7 produced better finished work in four of our small controlled tasks.
Gemini 3.8 Flash Limitations We Observed
The biggest issue was not a dramatic failure. It was extra correction work: a denominator mistake, vague status labeling, an awkward SEO sentence, and more tokens used on every completed task. T01 also ended before completing a tightly specified 180-word deliverable.
This test pack used text only. It does not evaluate image, audio, video, PDFs, computer use, true long-context retrieval, or a multi-hour agent run. Google’s model card also notes possible hallucinations, occasional slowness or timeouts, and higher token use at stronger effort levels.
Gemini 3.8 Flash Review: Final Verdict
Gemini 3.8 Flash is a meaningful official upgrade for agentic and long-horizon workloads, but it is not an automatic upgrade for every everyday task. It won one of six usable pairs in this review, while Gemini 3.7 Flash won four and used less time, fewer output tokens, and lower estimated cost.
The practical lesson is to test the work you actually ship. If your workload is long-running software engineering or document-heavy agent execution, Google’s official benchmark direction makes 3.8 worth a serious trial. If your priority is concise editing, small data analyses, or status-sensitive summaries, our first results say 3.7 remains highly competitive.
The same pattern appeared in our earlier Comparaison entre Gemini 3.6 Flash et Gemini 3.1 Pro: generation labels do not replace side-by-side output inspection. Run a recurring prompt in GlobalGPT, inspect the finished answer, and choose from evidence rather than the version number.
Gemini 3.8 Flash FAQ
When was Gemini 3.8 Flash released?
Google DeepMind published the Gemini 3.8 Flash model card on September 2, 2026.
Is Gemini 3.8 Flash generally available?
Yes. Google’s official product page lists Gemini 3.8 Flash as generally available.
What is the Gemini 3.8 Flash API model ID?
The official model ID is gemini-3.8-flash, which also matched the returned model ID in our API tests.
How much does Gemini 3.8 Flash cost?
Through December 31, 2026, Google lists $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, the rates rise to $1.50 and $7.50.
Is Gemini 3.8 Flash free?
Google’s pricing page lists a free tier with limited model access and free input and output tokens. Limits and model availability can differ by account and route.
What is the Gemini 3.8 Flash context window?
Gemini 3.8 Flash supports up to 1M input tokens and up to 64K tokens of text output, according to the official model card.
Does Gemini 3.8 Flash support images, audio, video, and PDFs?
Yes as inputs. Google’s product page lists text, image, video, audio, and PDF inputs; the model card lists text output.
Does Gemini 3.8 Flash support tools?
Google lists function calling, Search as a tool, and computer use. Tool availability can depend on the product and API route.
Is Gemini 3.8 Flash better than Gemini 3.7 Flash?
Not in every task. Gemini 3.8 Flash won one of six usable task pairs in this review, Gemini 3.7 Flash won four, and one pair tied. Google’s official evaluations emphasize stronger long-horizon agent and software-engineering performance.
Is Gemini 3.8 Flash good for coding?
Both models produced the correct minimal Python fix in our coding task. Google’s official positioning and DeepSWE claim make 3.8 especially relevant to longer software-engineering workflows, which our small bug fix did not measure.
Why did Gemini 3.8 Flash cost more in the tests if the published rate was the same?
Its recorded token usage was higher in every completed pair. A shared per-token rate can still produce a higher total cost when one response consumes more tokens.
Where can I access Gemini 3.8 Flash?
Google lists the Gemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google Antigravity. Availability can still vary by account, product, and region.
Are the benchmark results independently verified here?
No. The benchmark section attributes Google’s official claims directly. Artificial Analysis was unreachable during research, so its leaderboard data is not included as verified evidence in this draft.



