Reseña rápida del Gemini 3.8 Flash: precio, pruebas y veredicto

Reseña rápida del Gemini 3.8 Flash: precio, pruebas y veredicto

Model review · Official benchmarks · Same-prompt API tests

Gemini 3.8 Flash is Google’s new general-availability Flash model for long-horizon software engineering, agents, and complex knowledge work. The launch claims are strong. Our first controlled tests are much less tidy.

This Gemini 3.8 Flash review separates official facts from hands-on evidence. Google’s pages establish the release date, model status, context window, pricing, tools, and benchmark claims. Seven same-prompt API tasks show what the model actually returned beside Gemini 3.7 Flash.

If you want to repeat the comparison without managing separate model tabs, GlobalGPT provides one workspace for testing multiple AI models. For context on the previous generation, our Gemini 3.6: reseña rápida shows how quickly the Flash line has been changing.

Estado de la evidencia: six matched tasks produced usable answers from both models. One research-synthesis task ended mid-sentence for both and is excluded from the quality score. Artificial Analysis was unreachable during this review, so no third-party leaderboard number is presented as verified.
1–4Task wins: 3.8 vs 3.7
1Corbata de calidad
1Inconclusive task
6/6Completed pairs favored 3.7 on speed

Gemini 3.8 Flash Review: The Short Answer

Gemini 3.8 Flash did not beat Gemini 3.7 Flash in this first seven-task review. Across the six usable pairs, 3.8 won the frontend planning task, tied on a Python bug fix, and lost four tasks on status precision, data reasoning, bounded agent planning, and SEO editing. The remaining research task was inconclusive.

The gap was not about raw knowledge. It came from small defects in finished work: one wrong denominator claim, looser status labels, less economical prose, and a less concrete tool plan. Those are exactly the details that affect whether an answer can be used immediately.

Resumen de la reseña

MedidaGemini 3.8 FlashGemini 3.7 Flash
Tareas completadas14
Ties11
Completed-task elapsed time49.266s42.909s
Tokens de salida de tareas completadas9,1997,489
Estimated completed-task cost$0.036121$0.028569

These results describe six small text tasks on one API route and one configuration. They are not a universal ranking of either model.

What Is Gemini 3.8 Flash?

Google describes Gemini 3.8 Flash as its most intelligent Flash workhorse for coding and agents, with particular emphasis on long-horizon software engineering and knowledge work. It is based on Gemini 3.7 Flash rather than a completely separate architecture.

The positioning matters because this is not merely a cheaper general chatbot. Google is pitching 3.8 Flash as a production agent model that can reason across documents, tools, code, and longer task sequences while preserving Flash-level latency and scale.

Google DeepMind official Gemini 3.8 Flash product page
Google positions Gemini 3.8 Flash as its Flash model for complex agentic tasks at scale. Source: Google DeepMind.

Gemini 3.8 Flash Release Date

En ficha oficial del modelo was published on September 2, 2026. Google’s product page lists the model as generally available rather than preview-only.

What Changed From Gemini 3.7 Flash?

The official emphasis shifts toward longer-running agent tasks and finished artifacts. Google highlights software engineering, rigorous reasoning, tool execution, multimodal understanding, and complex enterprise work. In our small test pack, that positioning translated into the strongest frontend plan, but not more reliable finished answers across the board.

Gemini 3.8 Flash Specs and Availability

EspecificaciónGemini 3.8 FlashPruebas
ID del modelo de APIgemini-3.8-flashOfficial pricing page and returned API responses
EstadoGeneral availabilityGoogle DeepMind product page
Contexto de entradaHasta 1 millón de tokensFicha técnica oficial
Longitud máxima del texto64 000 tokensFicha técnica oficial
InputsText, images, audio, video, and PDFsPágina oficial del producto
SalidaTextoFicha técnica oficial
Soporte para herramientasFunction calling, Search as a tool, computer usePágina oficial del producto
Official access surfacesGemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google AntigravityFicha técnica oficial
Límite de conocimientosMarch 2026 for some domains; January 2025 for othersFicha técnica oficial

A 1M-token window is capacity, not proof of perfect long-document recall. The test pack below includes a short status-sensitive extraction task, but it does not establish million-token performance.

Gemini 3.8 Flash Pricing and API Cost

En Página oficial de precios de la API Gemini lists a launch rate through December 31, 2026 of $0,75 por cada millón de tokens de entrada y $3,75 por cada millón de tokens emitidos, including thinking tokens. On January 1, 2027, those rates rise to $1.50 and $7.50.

Published price change

Input · 2026
$0.75
Input · 2027
$1.50
Output · 2026
$3.75
Output · 2027
$7.50

Paid-tier price per 1M tokens. The January 2027 rates are double the introductory 2026 rates.

Gemini 3.8 Flash and Gemini 3.7 Flash had the same published introductory rate in our test period. Even so, 3.8 cost more in every completed pair because its recorded token usage was higher. Across T02–T07, the estimates were $0.036121 for 3.8 and $0.028569 for 3.7.

For a different point on the speed-and-cost curve, see our Gemini 3.5 Flash Lite review. A Lite model serves a different workload, so its lower price should not be treated as an automatic quality win.

Gemini 3.8 Flash Official Benchmarks

Google’s performance page emphasizes four areas: DeepSWE v1.1, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, and HLE-Verified. Its charts show Gemini 3.8 Flash ahead of Gemini 3.7 Flash on all three evaluations with exact published percentages.

Evaluación oficialGemini 3.8 FlashGemini 3.7 FlashLímite de las pruebas
DeepSWE v1.1Above 70% in Google’s chartLower than 3.8 in the chartGoogle says 3.8 leads in the efficient area; the public graphic does not expose a precise numeric label.
Vals Finance Agent v261.4%59.0%Official Google chart; the benchmark’s full methodology still matters.
Harvey’s Legal Agent Benchmark10.0%8.8%This benchmark result is not a guarantee of legal accuracy in deployment.
HLE-Verified54.9%53.6%It does not directly measure SEO editing, UI planning, or response economy.
Google DeepMind DeepSWE v1.1 benchmark chart for Gemini 3.8 Flash
Google’s official DeepSWE v1.1 chart places Gemini 3.8 Flash above 70% and in the chart’s most-efficient area. Source: Google DeepMind performance page.
Third-party benchmark note: Artificial Analysis could not be reached during research on September 4, 2026. This draft therefore uses official claims with clear attribution and leaves the independent leaderboard screenshot pending.

How We Tested Gemini 3.8 Flash

Both models received the same task prompt, temperature of 0, a 3,200-token output ceiling, and the same API route. T02 was rerun once for both models after the first 3.8 request failed technically; only the matched rerun is scored.

Campo registradoHow it was used
Returned model IDConfirmed that the requested model answered.
Real outputCompared correctness, completeness, status fidelity, and editing required.
Segundos transcurridosMeasured local end-to-end response time, not provider-side latency.
Tokens de entrada y salidaSupported task-level cost estimates at the official introductory rate.
Finish reasonSeparated usable answers from incomplete or budget-limited output.

Gemini 3.8 Flash Real-World Test Results

Each card keeps the task, the actual returned work, the measurements, and the editorial verdict together. The excerpts are shortened for readability but preserve the deciding details.

Test 1: Research Synthesis

Tarea: Produce a 180-word buying recommendation from a locked fact set, with three reasons, two unknowns, and a test plan.

Gemini 3.8 Flash

18.211s · 361 input / 3,196 output tokens · est. $0.012256

“We recommend selecting Model A… 3. Cost Parity: The introductory API price is the same for Model A and Model B”

Gemini 3.7 Flash

13.490s · 108 input / 3,196 output tokens · est. $0.012066

“We directly recommend piloting and purchasing Model A… The introductory API price remains the exact same for both models through December 3”

Both answers stop mid-sentence before the required unknowns and test plan. The 3.8 response reports basta, while 3.7 reports longitud; neither finish reason changes the usability judgment.

Verdict: Inconclusive. Neither model completed the requested deliverable.

Test 2: Python Bug Fix

Tarea: Return the three highest-value paid records, preserve input order on ties, handle empty input, explain the root cause, and add two tests.

Gemini 3.8 Flash

7.969s · 344 input / 1,517 output tokens · est. $0.005947

paid.sort(key=lambda r: r["amount"], reverse=True)
return paid[:3]

It explained Python’s stable sort and tested ranking, ties, filtering, and empty input.

Gemini 3.7 Flash

7.818s · 91 input / 1,359 output tokens · est. $0.005164

paid.sort(key=lambda r: r["amount"], reverse=True)
return paid[:3]

It returned the same minimal fix and covered the same acceptance criteria.

The code is functionally equivalent. The 3.7 answer is slightly shorter and cheaper, but that is not enough to separate output quality.

Verdict: Tie. Both produced a correct, usable patch.

Test 3: Textual Frontend Implementation Plan

Tarea: Plan five prioritized fixes, responsive behavior, and QA for a described checkout. No screenshot was supplied.

Gemini 3.8 Flash

10.128s · 340 input / 1,509 output tokens · est. $0.005914

It specified a 58/42 desktop grid, 48–52px CTA, nearby trust note, contextual discount alert, mobile order-summary accordion, focus states, and overflow checks.

Gemini 3.7 Flash

8.149s · 87 input / 1,416 output tokens · est. $0.005375

It proposed a similar grid and mobile stack, but introduced “256-bit Encrypted Checkout,” a detail absent from the task.

Both plans are implementable. The 3.8 plan has the stronger QA checklist and avoids inventing a specific encryption claim.

Verdict: Gemini 3.8 Flash. More complete and safer to hand to a frontend team.

Test 4: Status-Aware Extraction

Tarea: Summarize Orion and preserve confirmed, planned, approved, pending, unknown, and unverified states in a fact table.

Gemini 3.8 Flash

6.292s · 371 input / 1,122 output tokens · est. $0.004486

It labeled the second pilot “Confirmed / Pending” and wrote that the first pilot “successfully handled” 82% of cases.

Gemini 3.7 Flash

7.126s · 118 input / 1,152 output tokens · est. $0.004409

It kept the rollout “Planned (Unapproved),” legal review “Planned (Pending),” and production readiness “Unknown.”

Both retain the core facts, but 3.7 separates approval, planning, pending review, and unknown status more precisely. Both add “successfully” to the neutral 82% pilot statement.

Verdict: Gemini 3.7 Flash, slight win. Mayor fidelidad en la representación de los estados.

Test 5: Data Reasoning

Tarea: Calculate ROAS and paid conversion rate, handle a zero-spend denominator correctly, rank channels, and recommend a conditional $300 split.

Gemini 3.8 Flash

11.074s · 377 input / 2,355 output tokens · est. $0.009114

It correctly ranked Email 7.0×, Search 3.0×, Social 2.6×, and excluded Partner ROAS. It then said zero spend creates a zero denominator for both ROAS and CPA.

Gemini 3.7 Flash

9.106s · 124 input / 1,969 output tokens · est. $0.007477

It returned the same correct rankings and treated Partner ROAS as not measurable, then proposed $200 Email / $100 Search when email can scale.

The 3.8 answer makes a real mathematical error: CPA is spend divided by conversions, so $0 divided by 4 is $0, not a zero-denominator case.

Verdict: Gemini 3.7 Flash. Correct reasoning without the extra denominator mistake.

Test 6: Bounded Agentic Planning

Tarea: Use four named tools to collect one official API price, stop on unofficial or ambiguous evidence, and define completion in under 220 words.

Gemini 3.8 Flash

5.603s · 354 input / 948 output tokens · est. $0.003821

It mapped search, fetch, parse, and persist steps, then added authenticity, availability, and price-clarity gates.

Gemini 3.7 Flash

5.127s · 101 input / 894 output tokens · est. $0.003428

It used explicit tool-call syntax, named model_pricing.csv, required distinct input and output rates, and defined one verified row as completion.

Both plans respect the stop conditions. The 3.7 plan is more concrete about the available calls and the saved artifact without pretending the tools ran.

Verdict: Gemini 3.7 Flash, slight win. More execution-ready at lower response cost.

Test 7: SEO Writing and Editing

Tarea: Return HTML only, use “Gemini 3.8 Flash review” exactly once, preserve four locked facts, and avoid unsupported results, prices, dates, or access claims.

Gemini 3.8 Flash

8.200s · 379 input / 1,748 output tokens · est. $0.006839

<p>Our Gemini 3.8 Flash review evaluates the Google model by comparing official benchmark context with same-prompt API tests, with hands-on results pending as the evaluation records:</p>

Gemini 3.7 Flash

5.583s · 126 input / 699 output tokens · est. $0.002716

<p>This Gemini 3.8 Flash review examines the new Google model by comparing official benchmark context with same-prompt API tests, with hands-on results currently pending.</p>

Both satisfy the structure and locked facts. The 3.7 sentence is cleaner and reads like publishable SEO copy; 3.8’s “as the evaluation records” construction is awkward.

Verdict: Gemini 3.7 Flash. Better editorial finish with fewer tokens.

Gemini 3.8 Flash Test Scorecard

TareaGanadorDeciding difference
Síntesis de la investigaciónNo concluyenteBoth incomplete
Python bug fixCorbataEquivalent code and coverage
Frontend planGemini 3.8 FlashBetter QA and no invented encryption claim
Status extractionGemini 3.7 FlashMore precise state labels
Razonamiento basado en datosGemini 3.7 FlashNo CPA denominator error
Planificación agencialGemini 3.7 FlashMore concrete calls and artifact
Edición SEOGemini 3.7 FlashCleaner publishable sentence

The result is a useful counterweight to the launch story. Gemini 3.8 Flash may be stronger on the official agent and long-horizon evaluations Google chose to highlight, yet 3.7 produced better finished work in four of our small controlled tasks.

Gemini 3.8 Flash Limitations We Observed

The biggest issue was not a dramatic failure. It was extra correction work: a denominator mistake, vague status labeling, an awkward SEO sentence, and more tokens used on every completed task. T01 also ended before completing a tightly specified 180-word deliverable.

This test pack used text only. It does not evaluate image, audio, video, PDFs, computer use, true long-context retrieval, or a multi-hour agent run. Google’s model card also notes possible hallucinations, occasional slowness or timeouts, and higher token use at stronger effort levels.

Gemini 3.8 Flash Review: Final Verdict

Gemini 3.8 Flash is a meaningful official upgrade for agentic and long-horizon workloads, but it is not an automatic upgrade for every everyday task. It won one of six usable pairs in this review, while Gemini 3.7 Flash won four and used less time, fewer output tokens, and lower estimated cost.

The practical lesson is to test the work you actually ship. If your workload is long-running software engineering or document-heavy agent execution, Google’s official benchmark direction makes 3.8 worth a serious trial. If your priority is concise editing, small data analyses, or status-sensitive summaries, our first results say 3.7 remains highly competitive.

The same pattern appeared in our earlier Comparación entre Gemini 3.6 Flash y Gemini 3.1 Pro: generation labels do not replace side-by-side output inspection. Run a recurring prompt in GlobalGPT, inspect the finished answer, and choose from evidence rather than the version number.

Gemini 3.8 Flash FAQ

When was Gemini 3.8 Flash released?

Google DeepMind published the Gemini 3.8 Flash model card on September 2, 2026.

Is Gemini 3.8 Flash generally available?

Yes. Google’s official product page lists Gemini 3.8 Flash as generally available.

What is the Gemini 3.8 Flash API model ID?

The official model ID is gemini-3.8-flash, which also matched the returned model ID in our API tests.

How much does Gemini 3.8 Flash cost?

Through December 31, 2026, Google lists $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, the rates rise to $1.50 and $7.50.

Is Gemini 3.8 Flash free?

Google’s pricing page lists a free tier with limited model access and free input and output tokens. Limits and model availability can differ by account and route.

What is the Gemini 3.8 Flash context window?

Gemini 3.8 Flash supports up to 1M input tokens and up to 64K tokens of text output, according to the official model card.

Does Gemini 3.8 Flash support images, audio, video, and PDFs?

Yes as inputs. Google’s product page lists text, image, video, audio, and PDF inputs; the model card lists text output.

Does Gemini 3.8 Flash support tools?

Google lists function calling, Search as a tool, and computer use. Tool availability can depend on the product and API route.

Is Gemini 3.8 Flash better than Gemini 3.7 Flash?

Not in every task. Gemini 3.8 Flash won one of six usable task pairs in this review, Gemini 3.7 Flash won four, and one pair tied. Google’s official evaluations emphasize stronger long-horizon agent and software-engineering performance.

Is Gemini 3.8 Flash good for coding?

Both models produced the correct minimal Python fix in our coding task. Google’s official positioning and DeepSWE claim make 3.8 especially relevant to longer software-engineering workflows, which our small bug fix did not measure.

Why did Gemini 3.8 Flash cost more in the tests if the published rate was the same?

Its recorded token usage was higher in every completed pair. A shared per-token rate can still produce a higher total cost when one response consumes more tokens.

Where can I access Gemini 3.8 Flash?

Google lists the Gemini App, Gemini Enterprise Agent Platform, Google AI Studio, Gemini API, AI Mode, and Google Antigravity. Availability can still vary by account, product, and region.

Are the benchmark results independently verified here?

No. The benchmark section attributes Google’s official claims directly. Artificial Analysis was unreachable during research, so its leaderboard data is not included as verified evidence in this draft.

Comparte el post:

Entradas relacionadas