Claude 机型对比:Haiku 与 Sonnet 与 Opus,以及 Fable 的定位

claude-模型-比较-英雄

A practical 2026 guide to choosing by speed, API cost, context, task complexity, and the cost of a wrong answer.

For most everyday writing, analysis, coding, and tool-using work, start with Claude Sonnet 5. Choose Haiku 4.5 when low latency and high volume matter more than handling difficult edge cases. Choose Opus 5 when the task is complex, hard to verify, or expensive to get wrong. Fable 5.1 sits above Opus as a specialist escalation path, not the default.

Those recommendations are for general mixed workloads. Anthropic’s production guidance says most workloads should start with Opus 5, while its own task matrix separately positions Sonnet for everyday speed-and-capability work. The right starting point therefore depends on your risk tolerance and evaluation process, not just the model name.

GlobalGPT makes the broader workflow practical with one affordable subscription: Claude, GPT-5.6 Sol, Gemini 3 Pro, Kimi K3, and multiple AI functions share one dashboard for research, writing, analysis, coding, and refinement. Its CLI can also bring that model access into your terminal, development workflow, and existing production tools. Comparing outputs is one useful part of the workflow; for the broader family-level choice, see our ChatGPT 与 Claude 与 Gemini 的对比. The rest of this guide focuses on choosing within Claude.

Claude Models: The Quick Answer

The four current tiers solve different economic problems. The cheapest call is not always the cheapest completed task, and the strongest model is not automatically the best use of a budget.

Claude model decision guide

Choose by workload and failure cost

Use Sonnet for balanced everyday work, move down for efficiency, and move up when verification is difficult or mistakes are expensive.

俳句 4.5

Efficiency first

Start here for high-volume, latency-sensitive, clearly bounded tasks whose errors are cheap to detect and correct.

Move up whenEdge cases or reviewer correction time erase the token-price advantage.

第5首十四行诗

Everyday balance

Use for mixed writing, analysis, coding, visual understanding, and tool-using work that still receives normal review.

Move up whenLong-horizon reasoning, verification, or failure cost becomes the limiting factor.

作品第5号

Capability first

Start here for complex agentic coding, enterprise work, difficult diagnosis, or decisions that are costly to get wrong.

Move down whenA cheaper model repeatedly passes the same acceptance checks.

Fable 5.1

Exception path

Escalate only for demanding reasoning or long-running agent work when higher-effort Opus still fails your evaluation.

不要妄下结论A higher tier is automatically cheaper after quality, review time, and retries are counted.

The practical rule

Freeze one representative task and its acceptance checks. Choose the least expensive model that passes reliably, then count correction and verification time before calling it the cheapest option.

Source snapshot: Anthropic’s Choosing the right modelModels overview, verified September 8, 2026. Anthropic says most workloads start with Opus 5; its task matrix separately positions Sonnet 5 for everyday work and Haiku 4.5 for the lowest latency and price.

The Current Claude Model Lineup

As of September 8, 2026, Anthropic’s active lineup includes Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1. Haiku, Sonnet, and Opus are durable family names, but the version attached to each family changes over time. Pin the exact model ID you evaluated in production instead of assuming a family label will always resolve to the same snapshot.

Current modelOfficial speed labelAPI context最大输出Base API input / output
俳句 4.5最快200K64K$1 / $5 per MTok
第5首十四行诗快速1M128K$2 / $10 per MTok
作品第5号中度1M128K$5 / $25 per MTok
Fable 5.1较慢1M128K$10 / $50 per MTok
Anthropic current Claude model lineup with pricing, latency, context, and output limits
Anthropic’s model overview lists Haiku 4.5, Sonnet 5, Opus 5, and Fable 5.1 as the current lineup on September 8, 2026.

Claude.ai access is not the same as API specifications

The Claude.ai plan table describes consumer access and currently lists a 200K context window for individual plans. The API documentation lists model-specific context and output limits, including 1M context for Sonnet 5, Opus 5, and Fable 5.1. Do not use the API table to infer what a particular Claude.ai plan includes, or use the consumer plan table to infer an API model’s limits. See the Claude pricing and plan guide for the plan-level distinction.

Claude Haiku vs Sonnet vs Opus: Key Differences

The main trade-off is not a simple quality ranking. Moving up the lineup buys more capability for complex work, but it also increases base token cost and generally reduces speed. Context limits and reasoning controls differ too.

The practical trade-offs

Official labels and base API specifications, paired with practical starting points for common workloads.

尺寸俳句 4.5第5首十四行诗作品第5号
最佳起点Fast, repeated, bounded workMixed everyday knowledge workComplex, high-consequence work
Official speed label最快快速中度
API context / output200K / 64K1M / 128K1M / 128K
Reasoning control手动扩展思维Adaptive thinking and effortAdaptive thinking and effort
Base API price$1 input / $5 output$2 input / $10 output$5 输入/$25 输出

Where Fable fits: Fable 5.1 is officially labeled Slower, with 1M context, 128K maximum output, and $10 / $50 base input/output pricing per million tokens. Treat it as an escalation tier when Opus at higher effort still fails your evaluation.

速度和延迟

Anthropic labels Haiku 4.5 “Fastest,” Sonnet 5 “Fast,” Opus 5 “Moderate,” and Fable 5.1 “Slower.” These are qualitative product labels, not measured milliseconds or tokens per second. Actual latency depends on request size, reasoning settings, tools, load, and serving mode.

Context, output, and reasoning controls

Haiku’s smaller 200K API context and 64K maximum output are still substantial for many bounded jobs. Sonnet and Opus expand both limits to 1M and 128K. Haiku 4.5 uses manual extended thinking, while Sonnet 5 and Opus 5 support adaptive thinking and effort control. That difference is one reason a single identical reasoning setting cannot create perfectly equivalent tests across all models.

Which Claude Model Should You Use for Each Task?

工作量为什么升级时间
Classification, extraction, routing, support triage俳句 4.5Low latency and low base token cost suit clear, repeatable work.Edge cases or correction time erase the savings.
Writing, analysis, visual understanding, everyday coding第5首十四行诗Balanced speed, capability, context, and cost for mixed work.Long-horizon reasoning or verification becomes the bottleneck.
Complex coding, difficult diagnosis, autonomous agents作品第5号Capability matters more than the lowest per-token price.Higher-effort Opus still fails a demanding representative evaluation.
Exceptional long-running or reasoning-heavy workFable 5.1It is the current escalation path above Opus.Use only when the measured benefit justifies its higher cost.

These are starting points, not universal winners. For example, a short but high-stakes code change may justify Opus, while a long but highly structured extraction job may still fit Haiku. For practical workflow examples, see our guide to using Claude for coding.

When your shortlist extends beyond Claude, GlobalGPT lets you run the same representative prompt through Claude, GPT-5.6 Sol, Gemini 3 Pro, and Kimi K3 in one workspace. Keep the acceptance check constant: speed matters for repeated work, while answer quality and correction time determine whether a lower-cost result is actually cheaper.

Choose by the Cost of Failure, Not Just Task Type

Task labels hide the most important variable: what happens when the answer is wrong? The framework below is an editorial decision tool grounded in Anthropic’s efficiency-first and capability-first guidance. It is not an Anthropic benchmark.

Failure-cost ladder

Start lower when errors are cheap to catch; start higher when verification is incomplete or expensive.

Low cost

实例
Tagging, routing, draft cleanup, reversible summaries

开始
俳句 4.5

Escalate
When correction time removes the efficiency gain

Material

实例
Customer-facing drafts, reviewed analysis, tested production code

开始
第5首十四行诗

Escalate
When review cannot reliably catch important failures

昂贵

实例
High-autonomy changes, difficult diagnosis, consequential decisions

开始
作品第5号

Escalate
To Fable only after higher-effort Opus misses a frozen evaluation

When to Upgrade or Downgrade Your Claude Model

Move up only when a measured limit appears

Haiku to Sonnet

Upgrade when edge cases, instruction misses, or reviewer correction time exceed the savings.

Sonnet to Opus

Upgrade when long-horizon execution, complex reasoning, verification, or failure cost becomes the constraint.

Opus at higher effort

Test the demanding workload with frozen acceptance checks before moving to the top tier.

Opus to Fable

Escalate only when higher-effort Opus still fails the representative evaluation and the gain is worth the cost.

Downgrade rule: Move back down when a cheaper model repeatedly passes the same acceptance checks at the required latency and reliability.

Anthropic model selection matrix for Haiku, Sonnet, Opus, and Fable
Anthropic’s matrix maps four current Claude models to different capability, speed, and workload needs.

A Fixed Claude API Cost Example

For a fixed aggregate workload of 1 million input tokens plus 100,000 output tokens across multiple standard API requests, base token cost ranges from $1.50 with Haiku 4.5 to $15.00 with Fable 5.1. This is not one request, and it does not override per-request output limits.

One workload, four Claude models

The same token volume costs 2x, 5x, or 10x the Haiku total as you move up the current tiers.

工作量
1.0M input + 0.1M output
across multiple requests
模型Input calculationOutput calculation总计
俳句 4.51.0 x $10.1 x $5$1.50
第5首十四行诗1.0 x $20.1 x $10$3.00
作品第5号1.0 x $50.1 x $25$7.50
Fable 5.11.0 x $100.1 x $50$15.00

Included: Standard base input and output token rates from one Anthropic price snapshot.

Excluded: Prompt caching, Batch API, fast mode, geography, tool fees, taxes, discounts, and GlobalGPT credits.

USD base rates per million tokens from Anthropic’s Model pricing, verified September 8, 2026. Sonnet 5’s $2 input / $10 output rates are standard pricing.

The price table tells you the cost of tokens, not the cost of a completed result. A cheaper model can become more expensive if retries, corrections, and verification consume more staff time than the token savings.

Hands-On Claude Model Test: Same Prompts, Different Failure Modes

On September 9, 2026, we sent the same three prompts through the Anywhere API to Claude Haiku 4.5, Sonnet 5, Opus 5, and Fable 5.1. The tasks covered structured support triage, evidence-led writing, and auditable JavaScript debugging. We kept each model’s first valid output and evaluated it against rules frozen before the responses were reviewed.

The result was not a simple ranking. Sonnet, Opus, and Fable each finished on 26.5 out of 30, while their strengths and failures differed sharply. Haiku returned HTTP 503 on both allowed attempts for every task, so its output quality remains unknown in this test.

Test setup and overall result

Three frozen prompts, the same task-level token limits, no tools or search, and the first valid output retained for each model.

地点Anywhere API
日期September 9, 2026
任务Structured triage, launch note, JavaScript debugging
Retry ruleOnce for transport, 429/5xx, empty content, or malformed response
模型Completed tasksTotal quality scoreMost useful signal
克劳德俳句 4.50/3未计分All six permitted requests returned 503; route availability only.
克劳德·桑内特 53/326.5/30Only model to satisfy every launch-note requirement.
Claude Opus 53/326.5/30Strong debugging depth; missed a writing length constraint.
Claude Fable 5.13/326.5/30Clean structured outputs; launch note was truncated after one empty response.
The three scored models tied at 26.5/30, but for different reasons. This is a task-profile result, not a universal model ranking. Haiku is excluded from the tie because no output was available to score.

结构化输出

Test 1: Support-ticket triage

Task result: Fable 5.1, 10/10

Fable was the only model to return directly parseable JSON with all six labels and actions correct. Opus and Sonnet got the labels right but added Markdown fences; both also strengthened one supplied deadline, and Sonnet prematurely promised a correction or refund.

模型现状得分Anywhere latencyReported tokensWhat happened
克劳德俳句 4.5Route unavailable未计分不详不详Two 503 responses; no quality score.
克劳德·桑内特 5Output retained7.5/1010.39 s1005Correct labels, but fenced JSON and an invented refund/correction outcome.
Claude Opus 5Output retained8.5/1011.87 s1051Correct labels and strong actions, but fenced JSON and a strengthened deadline.
Claude Fable 5.1Output retained10/108.55 s680Directly parseable JSON; all frozen checks passed.
Full test input
You manage a SaaS support inbox. Classify every ticket below. Return JSON only with one top-level field named tickets. tickets must be an array in the original order, and every item must contain exactly ticket_id, category, urgency, and action. category must be one of billing, access, bug, feature. urgency must be one of low, medium, high. Use only the ticket text; do not invent policy outcomes or customer facts.

#481: We were charged twice for the same annual renewal. Payroll closes Friday and I need this investigated before then.
#482: Password reset says the email was sent, but nothing arrives. I checked spam and retried once.
#483: CSV exports round every currency amount to whole dollars, while the dashboard still shows cents.
#484: Please add a dark mode to the analytics dashboard.
#485: After our company domain changed this morning, every employee receives an SSO access denied message.
#486: Could you send me a copy of last month's invoice?
Claude Haiku 4.5: full operational result
No model output was produced. Both permitted Anywhere API attempts returned HTTP 503, so this route was marked unavailable and excluded from quality scoring.
Claude Sonnet 5: full output
```json
{
  "tickets": [
    {
      "ticket_id": "#481",
      "category": "billing",
      "urgency": "high",
      "action": "Investigate the duplicate annual renewal charge and issue a correction/refund before the customer's payroll closes Friday."
    },
    {
      "ticket_id": "#482",
      "category": "access",
      "urgency": "medium",
      "action": "Verify the account's email address and mail deliverability, then manually trigger or resend the password reset email."
    },
    {
      "ticket_id": "#483",
      "category": "bug",
      "urgency": "medium",
      "action": "Escalate to engineering to fix CSV export rounding so currency values match the cents-level precision shown on the dashboard."
    },
    {
      "ticket_id": "#484",
      "category": "feature",
      "urgency": "low",
      "action": "Log the dark mode request for the analytics dashboard and add it to the product feature backlog."
    },
    {
      "ticket_id": "#485",
      "category": "access",
      "urgency": "high",
      "action": "Immediately investigate SSO configuration tied to the changed company domain and restore access for all affected employees."
    },
    {
      "ticket_id": "#486",
      "category": "billing",
      "urgency": "low",
      "action": "Locate last month's invoice and send a copy to the customer."
    }
  ]
}
```
Claude Opus 5: full output
```json
{
  "tickets": [
    {
      "ticket_id": "481",
      "category": "billing",
      "urgency": "high",
      "action": "Escalate to billing to verify the duplicate annual renewal charge and confirm resolution before the customer's Friday payroll close."
    },
    {
      "ticket_id": "482",
      "category": "access",
      "urgency": "medium",
      "action": "Check email delivery logs for the reset message to that address and, if undelivered, trigger a verified reset path for the user."
    },
    {
      "ticket_id": "483",
      "category": "bug",
      "urgency": "medium",
      "action": "File a bug for CSV export currency rounding, noting the mismatch with dashboard values, and request reproduction steps or a sample export."
    },
    {
      "ticket_id": "484",
      "category": "feature",
      "urgency": "low",
      "action": "Log as a dark mode feature request for the analytics dashboard and acknowledge receipt without committing to a timeline."
    },
    {
      "ticket_id": "485",
      "category": "access",
      "urgency": "high",
      "action": "Escalate as an org-wide SSO outage tied to the domain change; have identity/auth review the SSO domain configuration immediately."
    },
    {
      "ticket_id": "486",
      "category": "billing",
      "urgency": "low",
      "action": "Send the requested prior-month invoice copy to the verified account contact."
    }
  ]
}
```
Claude Fable 5.1: full output
{"tickets":[{"ticket_id":"481","category":"billing","urgency":"high","action":"Investigate the reported duplicate annual renewal charge and respond before the customer's Friday payroll deadline."},{"ticket_id":"482","category":"access","urgency":"medium","action":"Check password reset email delivery logs for the user's address and resend or provide an alternate reset path."},{"ticket_id":"483","category":"bug","urgency":"medium","action":"Reproduce the CSV export rounding of currency amounts to whole dollars and escalate to engineering for a fix."},{"ticket_id":"484","category":"feature","urgency":"low","action":"Log the dark mode request for the analytics dashboard with the product team and acknowledge receipt."},{"ticket_id":"485","category":"access","urgency":"high","action":"Escalate the company-wide SSO access denied errors following the domain change and review SSO domain configuration."},{"ticket_id":"486","category":"billing","urgency":"low","action":"Locate and send the customer a copy of last month's invoice."}]}

First-valid-output policy. Latency is wall-clock time through Anywhere, not native Anthropic latency. Token counts are provider-reported; no direct API cost is inferred from this route.

Constraint following

Test 2: Evidence-led launch note

Task result: Sonnet 5, 10/10

Sonnet was the only model to satisfy every frozen requirement. Opus produced a strong note but exceeded the word limit and inferred that all non-respondents were enrolled students. Fable’s first response was empty; its allowed retry returned a visibly truncated note without the recommendation or caveat.

模型现状得分Anywhere latencyReported tokensWhat happened
克劳德俳句 4.5Route unavailable未计分不详不详Two 503 responses; no quality score.
克劳德·桑内特 5Output retained10/109.94 s796Met the word range, structure, evidence, recommendation, and caveat requirements.
Claude Opus 5Output retained8/109.49 s826Strong analysis, but 223 words under the frozen count method and one unsupported inference.
Claude Fable 5.1Output retained6.5/1019.27 s1685First response empty; retained retry was truncated before the recommendation and caveat.
Full test input
Write a 180-220 word internal launch note for a product director using only the facts below. Include: a subject line, a short opening paragraph, exactly three bullet points, a recommendation, and one clearly labeled caveat. Do not invent a budget resolution, percentages, statistical significance, or causal effects.

Facts:
- The pilot enrolled 120 students.
- At the 31 March checkpoint, 96 students had attended at least three sessions.
- A survey score averaged 4.3 out of 5 from 71 respondents.
- The program operations sheet lists a budget of $18,000.
- The finance approval note lists $20,000. It does not say whether the difference is a revision or a recording error.
- Participants were not randomly assigned and no baseline score was collected.
- The decision is whether to launch a second pilot, not whether to roll out permanently.
Claude Haiku 4.5: full operational result
No model output was produced. Both permitted Anywhere API attempts returned HTTP 503, so this route was marked unavailable and excluded from quality scoring.
Claude Sonnet 5: full output
**Subject: Pilot Program Results — Recommendation to Proceed with Second Pilot**

The initial pilot has concluded its first checkpoint period, and early engagement and satisfaction indicators are encouraging. Before committing to a second pilot, there are a few data points and open items the team should review together.

- Of the 120 students enrolled, 96 had attended at least three sessions as of the March 31 checkpoint.
- A post-session survey completed by 71 respondents produced an average score of 4.3 out of 5.
- Budget figures are inconsistent across documents: the program operations sheet lists $18,000, while the finance approval note lists $20,000; it is unclear whether this reflects a revision or a recording error.

**Recommendation:** Given the attendance and survey results, we recommend moving forward with a second pilot to gather more structured data — ideally with a defined comparison group and baseline measurement — before considering any broader rollout decision. This second pilot should also serve as an opportunity to reconcile the budget discrepancy with finance before funds are finalized.

**Caveat:** These results should be interpreted cautiously. Participants were not randomly assigned and no baseline score was collected, so we cannot attribute the survey results or attendance patterns to the program itself. This is a scoping decision for a second pilot, not a decision on permanent implementation.
Claude Opus 5: full output
**Subject: Pilot results and recommendation on a second pilot cohort**

The first pilot has completed its 31 March checkpoint, and we now have enough information to decide whether to run a second pilot. This note summarizes what we observed, what we recommend, and the main limitation on interpretation. No permanent rollout decision is on the table at this stage.

- **Enrollment and engagement:** 120 students enrolled. As of the 31 March checkpoint, 96 of them had attended at least three sessions.
- **Participant feedback:** The post-participation survey averaged 4.3 out of 5, based on 71 respondents. The remaining enrolled students did not respond, so the score reflects that subset only.
- **Budget discrepancy:** The program operations sheet lists a budget of $18,000, while the finance approval note lists $20,000. The finance note does not state whether the difference reflects a revision or a recording error, and we have not resolved which figure governs.

**Recommendation:** Proceed with planning a second pilot, conditional on Finance and Program Operations reconciling the two budget figures and confirming a single governing number before we commit spend. A second pilot should also collect baseline scores at intake.

**CAVEAT:** Participants were not randomly assigned and no baseline score was collected. The engagement and survey figures therefore describe what this group did and reported; they cannot tell us what the program caused or how it would perform with a different population.
Claude Fable 5.1: full output
**Subject:** Pilot checkpoint summary and second-pilot recommendation

Hi [Director],

Below is a summary of where the pilot stands at the 31 March checkpoint, along with a recommendation on whether to proceed to a second pilot. This note covers only the second-pilot decision, not permanent rollout.

- **Enrollment and attendance:** 120 students enrolled. At the 31 March checkpoint, 96 had attended at least three sessions.
- **Participant feedback:** A survey returned an average score of 4.3 out of 5 from 71 respondents.
- **Budget records:** The program operations sheet lists a budget of $18,000, while the finance approval note lists $20,000. The finance note does not indicate whether the difference is a revision or a recording error.

**

First-valid-output policy. Latency is wall-clock time through Anywhere, not native Anthropic latency. Token counts are provider-reported; no direct API cost is inferred from this route.

Code correctness

Test 3: Auditable JavaScript debugging

Task result: Opus 5 and Fable 5.1 tied, 10/10

All three available models found both defects and produced replacements that passed eight executable cases. Opus and Fable returned valid JSON and tied at 10/10. Sonnet’s code also passed, but the surrounding Markdown fence broke the direct-JSON requirement.

模型现状得分Anywhere latencyReported tokensWhat happened
克劳德俳句 4.5Route unavailable未计分不详不详Two 503 responses; no quality score.
克劳德·桑内特 5Output retained9/1013.17 s1471Replacement passed 8/8 cases; fenced output lost the JSON-format point.
Claude Opus 5Output retained10/1020.17 s1826Valid JSON, both defects identified, and 8/8 executable cases passed.
Claude Fable 5.1Output retained10/1015.89 s1409Valid JSON, concise diagnosis and 8/8 executable cases passed.
Full test input
Diagnose and repair this JavaScript function. Return JSON only with exactly four fields: root_cause (string), replacement (string containing the complete replacement function), tests (array of test-case strings), and remaining_risk (string). Do not claim that you executed tests. Make the smallest clear repair that satisfies the contract.

Contract:
- tasks is an array of zero-argument functions that return promises.
- Return the first fulfilled value in input order, even when that value is 0, false, or an empty string.
- Ignore rejected tasks.
- Return null when every task rejects or the array is empty.

Current code:
async function firstSuccessful(tasks) {
  const results = await Promise.all(tasks.map((task) => task()));
  return results.find(Boolean) || null;
}
Claude Haiku 4.5: full operational result
No model output was produced. Both permitted Anywhere API attempts returned HTTP 503, so this route was marked unavailable and excluded from quality scoring.
Claude Sonnet 5: full output
```json
{
  "root_cause": "Two bugs: (1) Promise.all rejects immediately if any task rejects, instead of ignoring rejected tasks; (2) results.find(Boolean) incorrectly treats falsy fulfilled values (0, false, '') as failures, causing them to be skipped in favor of null.",
  "replacement": "async function firstSuccessful(tasks) {\n  const settled = await Promise.allSettled(tasks.map((task) => task()));\n  const firstFulfilled = settled.find((r) => r.status === 'fulfilled');\n  return firstFulfilled ? firstFulfilled.value : null;\n}",
  "tests": [
    "test('returns null for empty array', async () => { expect(await firstSuccessful([])).toBeNull(); });",
    "test('returns null when all tasks reject', async () => { const tasks = [() => Promise.reject('err1'), () => Promise.reject('err2')]; expect(await firstSuccessful(tasks)).toBeNull(); });",
    "test('returns first fulfilled value even if falsy (0)', async () => { const tasks = [() => Promise.resolve(0), () => Promise.resolve(42)]; expect(await firstSuccessful(tasks)).toBe(0); });",
    "test('returns first fulfilled value even if false', async () => { const tasks = [() => Promise.resolve(false), () => Promise.resolve(true)]; expect(await firstSuccessful(tasks)).toBe(false); });",
    "test('returns first fulfilled value even if empty string', async () => { const tasks = [() => Promise.resolve(''), () => Promise.resolve('nonempty')]; expect(await firstSuccessful(tasks)).toBe(''); });",
    "test('ignores leading rejected tasks and returns first fulfilled in order', async () => { const tasks = [() => Promise.reject('err'), () => Promise.resolve('second'), () => Promise.resolve('third')]; expect(await firstSuccessful(tasks)).toBe('second'); });",
    "test('preserves input order even if later task resolves faster', async () => { const tasks = [() => new Promise((res) => setTimeout(() => res('slow-first'), 50)), () => new Promise((res) => setTimeout(() => res('fast-second'), 5))]; expect(await firstSuccessful(tasks)).toBe('slow-first'); });"
  ],
  "remaining_risk": "The fix assumes 'first fulfilled value in input order' means the first task by array index that fulfills, regardless of timing, which matches the contract wording but await Promise.allSettled waits for all tasks to complete (including slow ones) before returning, which may not be desired if early-exit performance is needed; also unhandled promise rejection warnings are avoided by allSettled, but if tasks throw synchronously instead of returning rejected promises, this is not explicitly tested here beyond the reject cases."
}
```
Claude Opus 5: full output
{
  "root_cause": "Two defects. (1) `Promise.all` rejects as soon as any task rejects, so a single rejected task makes the whole function reject instead of ignoring it and returning a later fulfilled value. (2) `results.find(Boolean) || null` selects by truthiness rather than by fulfillment, so legitimate fulfilled values such as 0, false, '', NaN, null, or undefined are skipped and/or coerced to null. Using `Promise.allSettled` plus a status check on the settled records fixes both, and returning `record.value` directly preserves falsy results.",
  "replacement": "async function firstSuccessful(tasks) {\n  const results = await Promise.allSettled(tasks.map((task) => task()));\n  const firstFulfilled = results.find((result) => result.status === 'fulfilled');\n  return firstFulfilled ? firstFulfilled.value : null;\n}",
  "tests": [
    "firstSuccessful([]) resolves to null (empty input).",
    "firstSuccessful([() => Promise.resolve('a'), () => Promise.resolve('b')]) resolves to 'a' (input order, not completion order).",
    "firstSuccessful([() => new Promise((res) => setTimeout(() => res('slow-first'), 50)), () => Promise.resolve('fast-second')]) resolves to 'slow-first' (order is by index, not by who settles first).",
    "firstSuccessful([() => Promise.reject(new Error('boom')), () => Promise.resolve('second')]) resolves to 'second' (rejection ignored, does not propagate).",
    "firstSuccessful([() => Promise.reject(new Error('a')), () => Promise.reject(new Error('b'))]) resolves to null (all reject).",
    "firstSuccessful([() => Promise.resolve(0)]) resolves to 0, not null (falsy number preserved).",
    "firstSuccessful([() => Promise.resolve(false)]) resolves to false (falsy boolean preserved).",
    "firstSuccessful([() => Promise.resolve('')]) resolves to '' (empty string preserved).",
    "firstSuccessful([() => Promise.resolve(undefined), () => Promise.resolve('later')]) resolves to undefined (first fulfilled wins even if its value is undefined).",
    "firstSuccessful([() => Promise.reject(new Error('x')), () => Promise.resolve(null), () => Promise.resolve('third')]) resolves to null because index 1 fulfilled with null — verifying null-as-value is not confused with the all-rejected sentinel by inspecting status.",
    "firstSuccessful([() => Promise.resolve(NaN)]) resolves to NaN (Number.isNaN on the result is true)."
  ],
  "remaining_risk": "The function still starts every task and awaits all of them, so there is no short-circuiting: a slow or never-settling later task delays or blocks the result, and side effects of later tasks always occur. A task that throws synchronously (rather than returning a rejected promise) will throw inside `tasks.map`, causing the returned promise to reject instead of being ignored; wrap the invocation in `Promise.resolve().then(task)` if that must also be tolerated. `null` is still returned both for all-rejected/empty input and when a task genuinely fulfills with `null`, so callers cannot distinguish those cases. Non-array or non-function inputs are not validated, and rejection reasons are discarded rather than aggregated (no AggregateError), which can hide diagnostics."
}
Claude Fable 5.1: full output
{"root_cause":"Two defects: (1) Promise.all rejects as soon as any task rejects, so rejected tasks are not ignored and the function throws instead of skipping them; (2) results.find(Boolean) skips legitimate falsy fulfilled values (0, false, '') and the trailing '|| null' also converts such values to null, violating the requirement to return the first fulfilled value even when it is falsy.","replacement":"async function firstSuccessful(tasks) {\n  const results = await Promise.allSettled(tasks.map((task) => task()));\n  const first = results.find((result) => result.status === 'fulfilled');\n  return first ? first.value : null;\n}","tests":["firstSuccessful([]) resolves to null","firstSuccessful([() => Promise.resolve('a'), () => Promise.resolve('b')]) resolves to 'a'","firstSuccessful([() => Promise.reject(new Error('x')), () => Promise.resolve('b')]) resolves to 'b' and does not throw","firstSuccessful([() => Promise.reject(new Error('x')), () => Promise.reject(new Error('y'))]) resolves to null","firstSuccessful([() => Promise.resolve(0), () => Promise.resolve(5)]) resolves to 0","firstSuccessful([() => Promise.resolve(false), () => Promise.resolve(true)]) resolves to false","firstSuccessful([() => Promise.resolve(''), () => Promise.resolve('z')]) resolves to ''","firstSuccessful([() => Promise.reject(new Error('x')), () => Promise.resolve(null)]) resolves to null (fulfilled null is returned as-is)","firstSuccessful([() => new Promise(r => setTimeout(() => r('slow'), 50)), () => Promise.resolve('fast')]) resolves to 'slow' (input order wins over settlement order)"],"remaining_risk":"All tasks are started concurrently and the function waits for every task to settle before returning, so a slow task delays the result even when an earlier task has already fulfilled; a sequential or early-exit strategy would be needed if latency matters. A task that throws synchronously (rather than returning a rejected promise) will still cause the function to reject, since the contract assumes tasks return promises. Promise.allSettled requires ES2020 or a polyfill."}

First-valid-output policy. Latency is wall-clock time through Anywhere, not native Anthropic latency. Token counts are provider-reported; no direct API cost is inferred from this route.

These Anywhere results show why model choice should follow the task instead of a single leaderboard. In GlobalGPT, you can apply the same comparison method to your own prompt across Claude and other leading models, then choose the output that best fits your quality bar, speed needs, and review workload.

How to reproduce the comparison fairly

  1. Freeze a representative task. Use the same intent, source material, tool access, and objective acceptance checks.
  2. Record the first valid output. Do not quietly give one model extra retries in the primary comparison.
  3. Separate quality from availability. Authentication, capacity, and route failures belong in an availability note, not the quality score.
  4. Measure the whole workflow. Count elapsed time, reviewer effort, correction time, and failed-task recovery.
  5. Choose the least expensive passing model. Re-test when prompts, tools, model versions, or failure costs change.

GlobalGPT 的应用场景

Anthropic guidance on when to choose Claude Fable 5.1 instead of Opus 5
Anthropic positions Fable 5.1 for demanding work when higher-effort Opus 5 evaluations still fall short.

For a deeper look at the upper tiers, continue to our advanced Claude model comparison or explore the dedicated Claude Fable 5.1 指南.

常见问题

What are the current Claude models?

As of September 8, 2026, Anthropic’s active lineup includes Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1. Exact versions and model IDs can change, so recheck the official model overview before deployment.

What is the difference between Claude Haiku, Sonnet, and Opus?

Haiku prioritizes speed and low base token cost, Sonnet balances speed and capability for everyday work, and Opus prioritizes capability for complex or high-consequence tasks. The right choice depends on your workload and verification cost, not an absolute ranking.

Which Claude model should I use for most tasks?

Start with Sonnet 5 for most everyday writing, analysis, coding, visual understanding, and tool-using work. Start with Opus 5 instead when failure is expensive or the work requires deeper reasoning and harder verification.

哪款 Claude 型号最适合编程?

Sonnet 5 is a practical starting point for everyday coding with normal review and tests. Use Opus 5 for difficult diagnosis, complex agentic coding, long-horizon execution, or changes that are costly to get wrong.

Which Claude model is best for writing?

Sonnet 5 is the balanced starting point for most drafting and editing. Haiku 4.5 can fit high-volume, tightly structured copy, while Opus 5 can be worth testing for difficult synthesis or high-stakes material that is hard to verify.

Is Claude Sonnet better than Opus?

Not universally. Sonnet 5 is faster and cheaper at base API rates, while Anthropic positions Opus 5 for more complex work. The better model is the least expensive option that reliably passes your own acceptance checks.

Is Claude Fable replacing Opus?

No. Anthropic positions Fable 5.1 as an escalation tier for demanding reasoning or long-running agent work when higher-effort Opus 5 still falls short, not as a universal replacement for Opus.

Do Claude.ai and the Claude API have the same context window?

No. Claude.ai plan limits and API model specifications are separate. Anthropic’s individual plan table currently lists a 200K context window, while the API documentation lists 1M context for Sonnet 5, Opus 5, and Fable 5.1.

When should I upgrade from Haiku to Sonnet or Opus?

Upgrade when failed acceptance checks, difficult edge cases, correction time, verification effort, or the cost of an error outweigh the lower model’s speed and token-price advantage. Move back down when a cheaper model repeatedly passes the same frozen checks.

How much do Claude Haiku, Sonnet, and Opus cost through the API?

At Anthropic’s base rates verified September 8, 2026, Haiku 4.5 costs $1 input and $5 output, Sonnet 5 costs $2 and $10, and Opus 5 costs $5 and $25 per million tokens. See Anthropic’s 官方定价页面 for current rates and additional pricing modes.

Final Verdict: Start with the Cheapest Model That Passes

For general mixed work, Sonnet 5 remains the most practical default, and it was the only model to pass every launch-note constraint in our test. Opus 5 and Fable 5.1 matched Sonnet’s 26.5/30 total through different strengths, while Haiku 4.5 could not be scored because the Anywhere route returned 503 for every permitted attempt. Treat these results as workload evidence, not a permanent universal ranking.

Your next step is simple: freeze one real task, define what a passing result looks like, run each serious candidate under comparable conditions, and include correction time in the final cost. The winning model is the least expensive one that passes reliably, not the one with the lowest token rate or the highest tier name.

GlobalGPT makes that method part of a broader, affordable AI workflow. Use Claude, GPT-5.6 Sol, Gemini 3 Pro, Kimi K3, and other leading models with multiple AI functions in one dashboard, then carry the work into your terminal, development workflow, and existing production tools through the GlobalGPT CLI. Side-by-side testing helps you choose; the larger advantage is keeping the complete workflow connected.

分享帖子:

相关帖子