Gemini 3.6 Flash 與 Gemini 3.1 Pro 比較:價格、效能測試及 6 項相同提示測試

Gemini 3.6 Flash 與 Gemini 3.1 Pro 比較

Gemini 3.6 Flash did not land quietly. Google released it as a faster, more efficient Flash model, but a lot of people had the same reaction: wait, where is the Pro upgrade?

That makes the comparison more interesting than a normal spec sheet. Gemini 3.6 Flash is supposed to be cheaper, faster, and stronger for coding, agents, and everyday multimodal work. Gemini 3.1 Pro、, meanwhile, still carries the “Pro” label and is positioned around deeper reasoning and harder tasks.

So we tested them the only way that actually matters: same task, same prompt, same input, same scoring rules.

In this review, we compare Gemini 3.6 Flash and Gemini 3.1 Pro across real tasks: research synthesis, coding repair, long-context summarization, data reasoning, screenshot analysis, and SEO writing. For each round, the main question is simple: which output would be easier to trust, edit, and reuse?

If you want to compare models without opening a pile of separate tabs, start from the GLBGPT model hub and run the same prompt across multiple AI models before trusting any launch claim. That matters for this kind of review, because the only honest way to compare models is to run the same work through them and look at the outputs.

API 狀態 Stable vs Preview

Gemini 3.6 Flash is listed as Stable; Gemini 3.1 Pro uses the Preview API label.

Official output price $7.50 vs $12+

Per 1M output tokens on standard paid API pricing, before long-prompt Pro tiers.

Our same-prompt run 4 wins vs 2

Flash won four task-level editorial calls; Pro won research synthesis and screenshot analysis.

Gemini 3.6 Flash vs Gemini 3.1 Pro Quick Answer

In our same-prompt API test, Gemini 3.6 Flash was the stronger overall practical model: it won four of the six task-level editorial calls and produced cleaner first drafts in several practical tasks.

Gemini 3.1 Pro Preview produced the more publishable SEO-style research recommendation and the richer screenshot audit. Flash answered the coding, long-context, data-reasoning, and SEO-editing tasks more effectively.

Where 3.6 Flash leads now First-draft usefulness

Flash had stronger results in coding, long-context summary, data reasoning, and SEO editing, while keeping the first draft easier to reuse.

Where 3.1 Pro can still win Clean reasoning and richer UI critique

Pro gave the cleaner data-reasoning response and the more detailed screenshot audit, especially when the answer needed richer critique.

What is not settled yet Longer and higher-stakes workloads

This run used six controlled API tasks. It does not prove every production workload, provider route, or long-context tier behaves the same way.

Gemini 3.6 Flash vs Gemini 3.1 Pro Official Specs

Google announced Gemini 3.6 Flash on July 21, 2026, alongside Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. In that launch post, Google frames 3.6 Flash as a more efficient Flash model for coding, knowledge work, multimodal performance, and agentic workflows.

The official Gemini API model page lists the model code as gemini-3.6-flash. It is marked Stable, supports text, image, video, audio, and PDF input, and returns text output. The listed token limits are 1,048,576 input tokens and 65,536 output tokens.

Gemini 3.1 Pro is a different kind of model. Google announced Gemini 3.1 Pro on February 19, 2026, positioning it around complex tasks and advanced reasoning. Google said the model was rolling out in preview to validate updates before general availability.

Google official Gemini 3.6 Flash release page showing the July 21, 2026 launch date.
Gemini 3.6 Flash announcement, captured July 22, 2026.
Google official Gemini 3.1 Pro release page showing the February 19, 2026 announcement.
Gemini 3.1 Pro announcement, captured July 22, 2026.

For API comparisons, the precise label matters. The official model page uses gemini-3.1-pro-preview, not a stable Pro label. It has the same listed input and output token limits as Gemini 3.6 Flash: 1,048,576 input tokens and 65,536 output tokens. To try the Pro side after reading the specs, open the Gemini 3.1 Pro model page and run one of the same prompts from this article.

If you need a setup walkthrough instead of raw model-code notes, the Gemini 3.1 Pro usage guide is the cleaner next step before you start testing prompts.

Gemini API model page showing gemini-3.6-flash as a stable model with a 1M input context window.
gemini-3.6-flash is listed as Stable with a 1M-token input window.
Gemini API model page showing gemini-3.1-pro-preview as a preview model.
gemini-3.1-pro-preview is listed as Preview with the same input/output limits.

Gemini 3.6 Flash vs Gemini 3.1 Pro Price Comparison

For standard paid Gemini API usage, Gemini 3.6 Flash is cheaper than Gemini 3.1 Pro Preview. The Gemini API pricing page lists Gemini 3.6 Flash at $1.50 per 1M input tokens and $7.50 per 1M output tokens.

Gemini 3.1 Pro Preview has two standard pricing tiers. For prompts up to 200K tokens, it is listed at $2.00 per 1M input tokens and $12.00 per 1M output tokens. For prompts above 200K tokens, it rises to $4.00 input and $18.00 output per 1M tokens. For a Pro-specific cost breakdown, keep the Gemini 3.1 Pro cost guide open beside the table.

模型標準投入價格標準輸出價格Important note
Gemini 3.6 Flash$1.50 / 1M tokens$7.50 / 1M tokensStable model; output price includes thinking tokens
雙子星 3.1 預覽$2.00 / 1M tokens under 200K prompt tokens$12.00 / 1M tokens under 200K prompt tokensPreview model; higher tier applies above 200K prompt tokens
Gemini 3.1 Pro Preview, long prompt$4.00 / 1M tokens above 200K prompt tokens$18.00 / 1M tokens above 200K prompt tokensImportant for long-context testing

Price alone does not decide model quality. A cheaper model that needs more retries can become expensive in practice. A slower model can still be worth it if it avoids a bad answer on a high-value task. If your prompts often push long context or quotas, check the Gemini 3.1 Pro limits guide before treating price as the whole decision. That is why our test records both cost and result quality for every prompt.

Gemini 3.6 Flash vs Gemini 3.1 Pro Benchmarks

Google 的 Gemini 3.6 Flash launch post gives the model a strong efficiency story. Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, improves on DeepSWE, improves on MLE Bench, and raises OSWorld-Verified performance compared with 3.5 Flash.

That is a good sign, but it does not answer the question we care about here. The real question is not only how 3.6 Flash compares with the earlier Gemini 3.5 Flash review; it is whether 3.6 Flash can now challenge Gemini 3.1 Pro in practical same-prompt work.

Gemini 3.1 Pro has a different official benchmark story. In the Gemini 3.1 Pro launch post, Google highlighted its advanced reasoning and said 3.1 Pro reached a verified ARC-AGI-2 score of 77.1%. That makes it easy to assume Pro should win the harder reasoning tests. But assumptions are exactly what this article is trying to avoid.

The third-party Artificial Analysis comparison we captured shows a much sharper contrast: Gemini 3.6 Flash appears faster, cheaper, and higher on its Intelligence Index than Gemini 3.1 Pro Preview. That is useful context, but it is not a Google official result, and it still does not replace same-prompt testing.

Artificial Analysis highlights for Gemini 3.6 Flash and Gemini 3.1 Pro Preview.
Artificial Analysis model comparison table. This screenshot supports the third-party Intelligence Index, blended price, output speed, time-to-first-token, and context-window values below.
來源公制Gemini 3.6 Flash雙子星 3.1 預覽How to use it
Google officialDeepSWE49%Not a direct 3.1 Pro comparison in the 3.6 Flash launch postOfficial claim mainly compares 3.6 Flash with 3.5 Flash
Google officialMLE Bench63.9%Not a direct 3.1 Pro comparison in the 3.6 Flash launch postUseful for coding/agentic expectations, not final proof
Google officialOSWorld-Verified83.0%Not a direct 3.1 Pro comparison in the 3.6 Flash launch postUseful for computer-use context
人工分析Intelligence Index5046Third-party benchmark context
人工分析輸出速度280 tokens/s119 tokens/sThird-party speed context
人工分析Time to first token11.71s33.44sThird-party latency context
人工分析Blended price$1.16 / 1M tokens$1.74 / 1M tokensUse separately from official API pricing

Data Figure: Artificial Analysis Snapshot

Third-party comparison captured July 22, 2026. Use this as benchmark context, not as a Google official claim.

Intelligence Index: 3.6 Flash
50
Intelligence Index: 3.1 Pro
46
Output speed: 3.6 Flash
280 t/s
Output speed: 3.1 Pro
119 t/s

How We Tested Gemini 3.6 Flash vs Gemini 3.1 Pro

Each test uses the same prompt, the same input, and the same scoring categories for both models. We do not treat one good answer as proof that a model is globally better. The goal is narrower and more useful: show how each model behaves across specific tasks.

We scored each output across these categories:

類別What we looked for
Result qualityIs the answer actually usable?
Instruction followingDid the model obey format, limits, and constraints?
精確度Did it avoid invented facts, prices, links, or claims?
推理深度Did it catch tradeoffs and hidden constraints?
結構Is the output easy to scan and reuse?
速度End-to-end response time in seconds
成本Estimated from official API input/output token pricing

The test was run through the same API route using the exact requested model IDs: gemini-3.6-flashgemini-3.1-pro-preview. The raw outputs, token usage, latency, and official API cost estimates were logged for every task.

If coding is the main reason you are comparing these two models, pair this result with the 最適合編碼的 AI 模型 guide instead of judging by one bug-fix prompt alone.

Method note: this is a same-prompt API workflow test, not an industry benchmark. Latency means local end-to-end elapsed time from this workspace. Cost is estimated from official Gemini API Standard paid-tier prices checked on July 22, 2026. The screenshot task used a synthetic checkout/dashboard image created for this test.

If you are still choosing a default model for broader work, the best AI models guide is a better next step than reading this Gemini pair in isolation.

And if your real decision is Google versus OpenAI rather than Flash versus Pro, the GPT-5 vs Gemini 2.5 Pro comparison is a better next step than another Gemini-only benchmark table.

T01 Research synthesis

Pro’s recommendation read more like polished SEO-team copy, although one matrix note misstated the source pack.

專業
T02 Coding bug fix

Both solved it, but Flash had more robust decimal assertions in the tests.

閃光燈
T03 Long-context summary

Pro was richer, but Flash was cleaner and safer on preview-status wording.

Flash, slight
T04 Data reasoning

Both did correct math; Flash used the more familiar division notation and a cleaner scan-friendly structure.

閃光燈
T05 Screenshot analysis

Flash was accurate; Pro gave richer trust, typography, and mobile-risk detail.

Pro, slight
T06 SEO writing/editing

Flash sounded more natural and clean; Pro was safe but too dramatic in tone.

閃光燈

Gemini 3.6 Flash vs Gemini 3.1 Pro Same-Prompt Results

Instead of hiding the test data in one combined scorecard, this section shows each prompt as its own Gemini 3.6 Flash vs Gemini 3.1 Pro comparison card. Each card includes same-prompt API run data plus short excerpts from the actual raw outputs, so the comparison is not just a score table.

Task Result / T01

研究綜述

Winner: Gemini 3.1 Pro Preview
Gemini 3.6 Flash
得分
4.23
Elapsed
22.8s
成本
$0.0471
輸出
6,206 tokens
雙子星 3.1 預覽
得分
3.76
Elapsed
48.0s
成本
$0.0874
輸出
7,217 tokens

Task setup: Each model had to write a 120-160 word internal buying recommendation from a controlled source pack only. The pack included model IDs, official API prices, buyer constraints, unknown evidence, and a strict four-item missing-evidence list.

Side-by-side screenshot of Gemini 3.6 Flash and Gemini 3.1 Pro Preview research synthesis outputs.

Actual Flash output
For our five-person SEO content team, we recommend setting gemini-3.6-flash as the primary default model and using gemini-3.1-pro-preview selectively.
Actual Pro output
(Pricing for prompts > 200k tokens is not in source pack).

What this shows: Pro organized the recommendation in a more publishable SEO style and used a more measured buying tone. Flash was more precise on the supplied pricing tiers; Pro’s matrix incorrectly said one Flash pricing detail was missing from the source pack.

Editorial verdict: Pro wins this writing-oriented task because its recommendation sounds closer to a finished internal SEO brief, but its pricing note should be corrected before publication.

Task Result / T02

Coding bug fix

Winner: Gemini 3.6 Flash
Gemini 3.6 Flash
得分
4.73
Elapsed
9.1s
成本
$0.0166
輸出
2,131 tokens
雙子星 3.1 預覽
得分
4.51
Elapsed
13.2s
成本
$0.0231
輸出
1,853 tokens

Task setup: The input was a small JavaScript cost calculator with two seeded bugs: it treated per-million-token prices as per-token prices, and it used the higher Pro tier at >=200,000 tokens instead of only above 200,000.

Actual Flash output
console.assert(Math.abs(actual – expected) < 1e-9, `Test 1 Failed: Expected ${expected}, got ${actual}`);
Actual Pro output
console.assert(test1 === 0.00525, `Test 1 Failed: Expected 0.00525, got ${test1}`);

What this shows: Both models fixed the module with a minimal patch, but Flash wrote safer tolerance-based floating-point tests while Pro used strict decimal equality.

Editorial verdict: Both answers worked. Flash was more production-ready because its tests used safer decimal assertions.

Task Result / T03

Long-context summary

Winner: Gemini 3.6 Flash, slight
Gemini 3.6 Flash
得分
4.53
Elapsed
15.0s
成本
$0.0455
輸出
3,489 tokens
雙子星 3.1 預覽
得分
4.2
Elapsed
28.2 秒
成本
$0.0732
輸出
3,949 tokens

Task setup: The prompt used a repeated planning transcript with late conflict notes about missing screenshots, token-cost reporting, Friday publishing pressure, model-id safety, and API failure handling.

Actual Flash output
Confirm Gemini 3.1 Pro is not described as "stable" if its model ID is "preview".
Actual Pro output
Gemini 3.1 Pro is confirmed stable, or the copy is adjusted to avoid calling it stable if the ID says "preview".

What this shows: Flash preserved the safer editorial rule: do not call Gemini 3.1 Pro stable when the API model ID says preview. Pro was richer, but its checklist wording could be read as confirming stability first.

Editorial verdict: Pro gave more texture, while Flash was cleaner and safer on the preview-status wording.

Task Result / T04

Data reasoning

Winner: Gemini 3.6 Flash
Gemini 3.6 Flash
得分
4.18
Elapsed
10.5s
成本
$0.0181
輸出
2,368 tokens
雙子星 3.1 預覽
得分
4.21
Elapsed
17.1s
成本
$0.0276
輸出
2,258 tokens

Task setup: The input was a 14-day channel table with spend, clicks, signups, paid conversions, revenue, and time on page. The model had to rank ROAS, handle the zero-spend channel separately, rank paid conversion rate, and recommend a next $300 budget split.

Side-by-side screenshot of Gemini 3.6 Flash and Gemini 3.1 Pro Preview data reasoning outputs.

Actual Flash output
Organic LinkedIn ($45 revenue on $0 spend; math results in undefined/infinite ROAS).
Actual Pro output
Signups are a vanity metric here because they do not guarantee actual revenue.

What this shows: Both models reached the correct rankings and budget split. Flash made the arithmetic easier to scan by using the familiar division sign instead of slash notation, and its numbered layout stayed consistent across both calculations.

Next step: For spreadsheet-style ranking, budget splits, and channel analysis, compare this task with the 最適合用於資料分析的人工智慧工具 指南。.

Editorial verdict: Flash wins this task on presentation. Pro gave a useful explanation of signup quality, but Flash’s calculation format is quicker to read and verify.

Task Result / T05

Screenshot analysis

Winner: Gemini 3.1 Pro Preview, slight
Gemini 3.6 Flash
得分
4.68
Elapsed
11.2s
成本
$0.0165
輸出
1,950 tokens
雙子星 3.1 預覽
得分
4.57
Elapsed
20.7s
成本
$0.0296
輸出
2,261 tokens

Task setup: The input image was a synthetic GLBGPT checkout/dashboard screenshot created for this test. It showed a Pro monthly checkout, a discount warning, a Pay now button, an order summary, delayed tax text, and a mobile trust-badge note.

Synthetic GLBGPT checkout dashboard screenshot used for the T05 screenshot analysis test.
Actual Flash output
Layout sparseness affecting visual trust: the excess empty white space inside both main panels creates an unpolished aesthetic.
Actual Pro output
Poor button typography: the "Pay now" text inside the green button is disproportionately small and uses a thin, black font.

What this shows: Pro identified more design-specific issues, including button typography, pricing trust, and mobile trust-signal placement. Flash’s audit was accurate but less specific.

Editorial verdict: Pro wins narrowly because its observations give a designer more concrete fixes to make.

Task Result / T06

SEO writing/editing

Winner: Gemini 3.6 Flash
Gemini 3.6 Flash
得分
4.39
Elapsed
19.9s
成本
$0.0366
輸出
4,818 tokens
雙子星 3.1 預覽
得分
4.11
Elapsed
34.9s
成本
$0.0641
輸出
5,291 tokens

Task setup: Each model had to rewrite a GLBGPT comparison section as HTML only, preserve locked commercial claims exactly, include the exact CTA URL, avoid invented test results, and stay between 180 and 260 words.

Gemini 3.6 Flash vs Gemini 3.1 Pro: Final Verdict

Gemini 3.6 Flash finishes ahead in this comparison, winning four of the six same-prompt tasks. Its strongest results came from coding, long-context summarization, data reasoning, and SEO editing. Gemini 3.1 Pro Preview still produced the better research synthesis and the more detailed screenshot audit, showing that the Pro model can deliver a more polished or observant answer when the task rewards depth.

Flash was also faster and cheaper across this controlled API run. That combination made it the stronger overall result here, but not a universal replacement for Pro: the actual outputs show that model quality still changes with the task, even when the prompt stays exactly the same.

For your own Gemini 3.6 Flash 與 Gemini 3.1 Pro 比較 comparison, run the same real prompt through both models in GLBGPT and judge the finished work side by side.

Gemini 3.6 Flash vs Gemini 3.1 Pro FAQ

When was Gemini 3.6 Flash released?

Google announced Gemini 3.6 Flash on July 21, 2026. The official Gemini API model page lists gemini-3.6-flash as a stable model.

Is Gemini 3.1 Pro the same as Gemini 3.1 Pro Preview?

For API pricing and technical comparison, use the official label gemini-3.1-pro-preview. Google’s February 2026 launch post says Gemini 3.1 Pro was released in preview to validate updates before general availability.

When will Gemini 3.5 Pro or Gemini 3.6 Pro be released?

As of July 22, 2026, Google has made an official statement about Gemini 3.5 Pro, but not a dated launch promise. In the Gemini 3.6 Flash announcement, Google says Gemini 3.5 Pro is currently testing with partners and that it plans to make the model broadly available “as soon as it’s ready.” I did not find an official Google release date or official announcement for Gemini 3.6 Pro.

Does Gemini 3.6 Flash replace Gemini 3.1 Pro?

Not officially. Gemini 3.6 Flash is a stable Flash model focused on speed, cost, coding, multimodal work, and agentic execution. Gemini 3.1 Pro Preview is positioned around more advanced reasoning and complex tasks. The same-prompt tests in this article are designed to show where the practical gap is now.

Can public benchmarks decide the winner?

No. Benchmarks help set expectations, but they do not replace controlled same-prompt testing. This article separates official benchmark claims, third-party benchmark data, and our own task-level test results.

Why does this article use Gemini 3.1 Pro Preview instead of just Gemini 3.1 Pro?

The article can use “Gemini 3.1 Pro” in reader-facing prose, but the technical tables should use gemini-3.1-pro-preview. That is the official API label used in the Gemini model and pricing pages.

Is Gemini 3.6 Flash cheaper than Gemini 3.1 Pro Preview?

Yes, under standard paid Gemini API pricing. Gemini 3.6 Flash is listed at $1.50 per 1M input tokens and $7.50 per 1M output tokens. Gemini 3.1 Pro Preview starts at $2.00 input and $12.00 output per 1M tokens, with a higher tier for prompts above 200K tokens.

Do both models support a 1M-token context window?

The official Gemini API model pages list 1,048,576 input tokens for both gemini-3.6-flashgemini-3.1-pro-preview. In practice, the long-context test still matters because a large window does not guarantee perfect recall.

Should the article include official screenshots?

Yes, but screenshots should support the claims rather than carry the whole article. The main comparison should be shown with readable tables, data cards, price bars, benchmark summaries, and same-prompt result tables.

What tasks were included in the same-prompt test?

The API run included six tasks: research synthesis, coding bug fix, long-context summary, data reasoning, screenshot analysis, and SEO writing/editing. Each run recorded latency, input tokens, output tokens, estimated cost, and a judge note.

Can this article say Gemini 3.6 Flash is better than Gemini 3.1 Pro?

It can say Gemini 3.6 Flash was better overall in this specific same-prompt API test. It should not claim that Flash is globally better for every workload, because Pro produced stronger results in research synthesis and screenshot analysis.

Where were the same-prompt tests run?

The tests were run through the same API route using the exact model IDs gemini-3.6-flashgemini-3.1-pro-preview. This is not a GLBGPT UI test, and the local timing should be read as end-to-end request time rather than provider-side latency.

Why did Gemini 3.6 Flash win overall if Pro won some quality categories?

Flash won because it produced the stronger reusable answer in more tasks. Pro’s research synthesis and screenshot analysis were stronger, but Flash’s wins covered coding, long-context summary, data reasoning, and SEO editing.

Why mention GLBGPT in the comparison?

Model comparisons are easier when readers can test multiple models in one place. GLBGPT is relevant here because the reader’s next action is not just reading specs; it is running the same prompt across models and comparing the outputs.

分享文章:

相關文章