Grok 4.7 レビュー:価格、ベンチマーク、API アクセス、および変更点

一言で言えば

Grok 4.7 is a serious coding and knowledge-work upgrade, but the evidence is still early: official launch claims and independent benchmarks point to improvement, not a universal win.

簡単な回答

Grok 4.7 is xAI’s September 21, 2026 upgrade aimed at coding and difficult knowledge-work tasks. xAI says it spends longer on hard problems, checks its work more carefully, and adds a new safeguard stack while keeping the same price and speed as Grok 4.6. The launch page also presents it as twice as fast and half the price of comparable models, a vendor comparison that needs to be read alongside the underlying task and route. Independent Artificial Analysis data gives it a small overall lift over Grok 4.6 and records a 500K context window. In our three-task API test, Grok 4.7 was strongest at debugging and source-boundary discipline; its project plan preserved uncertainty well but was less actionable than requested. That is useful hands-on evidence, not proof that it beats every competing model.

What is Grok 4.7?

xAI Introducing Grok 4.7 headline and official positioning
xAI positions Grok 4.7 around coding and knowledge work, with speed and price claims stated on the launch page.

Grok 4.7 is the latest model in xAI’s Grok line, announced on September 21, 2026. The official xAI announcement positions it as the company’s most capable model for coding and knowledge work. The meaningful change is not a new consumer-facing feature list. It is a change in how the model is expected to spend effort: more time on difficult tasks, more self-checking, and a safety layer intended to make that extra capability more controllable.

That positioning matters because it describes a workflow model rather than a chatbot personality. Grok 4.7 is being sold as a model for tasks that benefit from sustained analysis: debugging, planning, technical writing, research synthesis, and long prompts that need the model to keep several constraints in view.

The boundary is equally important. xAI’s announcement is a vendor statement about intended behavior. It is not an independent proof that Grok 4.7 beats every competing model, every prompt, or every coding repository. Treat the launch page as the source for what xAI claims, then use independent measurements and your own task to decide whether those claims matter.

Grok 4.7 vs Grok 4.6: what actually changed?

SpaceXAI official X post sharing the Grok 4.7 launch comparison graphic
SpaceXAI repeats the launch framing on X; this is official positioning, not independent validation.

The most defensible summary is an effort and reliability upgrade, not a completely different product.

エリアGrok 4.7What the evidence supports
リリースSeptember 21, 2026Date stated on the official xAI launch page
Main positioningCoding and knowledge workOfficial xAI positioning
Difficult tasksLonger effort and more careful self-checkingOfficial xAI description; our three-task API run offers a small hands-on check, not a matched comparison
安全New safeguard stackOfficial xAI description
Price and speedSame as Grok 4.6, according to xAIA launch claim, not a third-party price audit
文脈500K tokens in the Artificial Analysis profileIndependent model profile; verify the route you use
Reasoning controlsLow through xhigh in Artificial Analysis’ listingIndependent route/model metadata

This is a useful distinction for buyers. If Grok 4.6 already fit your workflow, 4.7 is a low-friction upgrade to evaluate. If 4.6 did not fit because of ecosystem, latency, reliability, or access constraints, a higher benchmark number alone does not solve those problems.

For a broader model-buying framework, the GPT-5.6 価格ガイド そして GPT-5.6の比較 are useful reminders to separate model quality from plan economics and workflow fit.

Capabilities: where Grok 4.7 should help

コーディングとデバッグ

xAI’s strongest launch framing is coding. The promise is not simply that Grok 4.7 writes more code; it is that it can stay with a hard engineering problem longer, inspect its own answer, and return a more considered patch or explanation. That is the behavior to look for in a real repository: does it preserve the existing design, identify the root cause, add useful tests, and state what it still cannot know?

Do not evaluate the model with a greenfield “build a todo app” prompt. Give it a bounded bug with a failing test, a small set of files, and explicit constraints. Ask for a diagnosis before a patch. This makes it easier to see whether the longer reasoning claim produces less review work or merely longer prose.

The same rule applies to teams comparing different coding routes. A DeepSeek V4 Pro レビュー そして GLMコーディング計画のレビュー provide useful comparison context, but neither should be treated as a substitute for testing the repository and toolchain you actually use.

Long-context knowledge work

Artificial Analysis lists a 500K-token context window for Grok 4.7. That is large enough to change how you approach source-heavy work: a long technical specification, a set of meeting notes, a codebase slice, or a research packet can stay in one conversation instead of being repeatedly summarized.

Context capacity is not the same as context quality. A model can accept a large input and still miss a constraint near the middle, over-weight the latest paragraph, or produce a confident summary that needs source checking. The practical test is retrieval under pressure: place a few decisive facts in different parts of a long document, ask for a constrained decision, and check whether the answer cites the right passages.

SpaceXAI’s official X account repeated the launch framing on September 22, calling Grok 4.7 a notable improvement over Grok 4.6 at the same price and speed and sharing the model comparison graphic. That post is useful evidence of the company’s public positioning, not independent validation or a community consensus.

If the model is used for research, keep the evidence loop visible. Our Gemini と Perplexity の比較 そして Perplexityの代替品ガイド make the same practical point: a polished answer is not a source audit.

Self-checking and safety

xAI says Grok 4.7 checks its work more carefully and introduces a new safeguard stack. Those are valuable goals, especially when the model is asked to take a long chain of decisions. They are not a guarantee that the output is correct or safe for automatic execution.

For production use, keep ordinary controls around the model: tests, permissions, human review, citation checks, and a clear stop condition. The model’s self-check is another signal to inspect, not a replacement for the controls around it.

Grok 4.7 benchmarks: what the numbers prove

xAI Grok 4.7 launch page with benchmark, safety, and pricing sections
Scroll inside the frame to inspect the full xAI launch page; the surrounding article stays compact.
Artificial Analysis Grok 4.7 benchmark article with Intelligence Index and AA-Briefcase takeaways
Artificial Analysis reports the independent Intelligence Index, AA-Briefcase, and coding-agent measurements.

The most useful independent snapshot currently comes from Artificial Analysis’ Grok 4.7 benchmarking article, dated September 21, 2026.

測定報告された結果読み方
人工知能分析指数46Two points higher than Grok 4.6 in the same index
AA-ブリーフケース1,657 Elo111 Elo above Grok 4.6 in that evaluation
コンテキストウィンドウ50万トークンA model-profile field, not proof of perfect long-context retrieval
推論努力Low to xhighA route-level control listed by Artificial Analysis

The direction is encouraging: the independent index shows improvement over the previous version, and the larger Elo gap in AA-Briefcase suggests a more noticeable gain on that particular evaluation. But the table still describes two datasets, not a universal ranking. Benchmark composition, tool access, sampling, and model settings can change the result.

The official xAI launch page shows Grok 4.7 at 46.3% on CursorBench 4.0 in its comparison table, alongside Grok 4.6 at 40.4%, GPT-5.6 Sol at 41.7%, and Fable 5.1 at 51.8%. The page also shows 71.0% on DeepSWE v1.1 at high effort, 64.0% on EEBench, and 1,657 on AA-Briefcase v1.1. These are vendor-presented comparison figures, so treat them as official reported results rather than a substitute for independent replication. The safer conclusion is that xAI is targeting serious coding and professional knowledge work, while independent measurements show a meaningful improvement on some tasks and flatter results on others.

Hands-on Grok 4.7 test: three practical tasks

We ran three fixed prompts against the grok-4.7 model ID through an OpenAI-compatible test API on September 23, 2026. This was a single-route evaluation, not the direct xAI API and not a matched competitor test. The visible answer from each run is preserved below; private reasoning traces are excluded.

タスク観測時刻結果
TypeScript concurrency bug333.7 seconds, streamedStrong pass: correct shared-promise diagnosis, minimal patch, two useful tests, and a careful limit on what the snippet could prove
Conflicting research notes68.7 secondsStrong pass: preserved attribution, calculated both allowance scenarios, and refused approval until the source conflict was verified
Messy one-week plan50.6 secondsPartial pass: preserved deadlines, contradictions, unknown owners, and five risks, but did not turn the work into a clear day-by-day schedule

The latency numbers describe this route and these prompts only. The coding task first hit a non-streaming gateway timeout, then completed with streaming enabled. That is a practical deployment warning about transport and timeout handling, not a general Grok 4.7 speed score.

Test 1: TypeScript concurrency debugging

Test prompt
Task 1 — TypeScript API handler concurrency debugging

A small TypeScript handler intermittently returns HTTP 200 with an empty body when two requests arrive close together.

Files provided:

src/cache.ts
```ts
let current: Promise<string> | null = null;

export async function getPayload(load: () => Promise<string>): Promise<string> {
  if (!current) {
    current = load();
  }
  const value = await current;
  current = null;
  return value;
}
```

src/handler.ts
```ts
import { getPayload } from './cache';

export async function handler(load: () => Promise<string>) {
  const payload = await getPayload(load);
  return { status: 200, body: payload };
}
```

The production symptom is intermittent only when two requests overlap. The loader itself is safe to call concurrently and returns a different payload for each request. No other files are available.

Return exactly these sections:
1. Root cause
2. Smallest safe patch (show a unified diff)
3. Two regression tests (TypeScript or Jest-style code)
4. Assumptions and limits
Do not claim to have run the code.
Grok 4.7 answer
1. Root cause

`current` is module-level state shared by every call. On one JavaScript turn the null check and assignment are atomic, so two overlapping calls do not both run `load()`. The first call stores its promise; a second call that arrives before that promise settles skips its own `load()` and awaits the same promise. `handler` then returns HTTP 200 whose `body` is the first call’s string, not the second request’s payload. That window exists only while a call is in flight, so the failure shows up only when requests overlap.

`current = null` runs in each continuation after `await` has already captured the promise. It does not make the other waiter re-read the variable, so these files alone do not resolve that waiter to `null` or `""`. What they do produce on overlap is a shared resolution and a skipped per-request load. The clear also makes the slot a bad single-flight cache: the first continuation to resume drops `current`, and a later caller can start a different `load()` while another caller is still unwinding.

2. Smallest safe patch (show a unified diff)

Each request must call `load()` itself. Concurrent `load()` calls are allowed, and each returns a different payload, so deduping on one shared promise is unsafe. `handler.ts` does not need a change.

```diff
--- a/src/cache.ts
+++ b/src/cache.ts
@@ -1,10 +1,3 @@
-let current: Promise<string> | null = null;
-
 export async function getPayload(load: () => Promise<string>): Promise<string> {
-  if (!current) {
-    current = load();
-  }
-  const value = await current;
-  current = null;
-  return value;
+  return await load();
 }
```

3. Two regression tests (TypeScript or Jest-style code)

```ts
import { handler } from '../src/handler';

function deferred<T>() {
  let resolve!: (value: T) => void;
  let reject!: (reason: unknown) => void;
  const promise = new Promise<T>((res, rej) => {
    resolve = res;
    reject = rej;
  });
  return { promise, resolve, reject };
}

test('overlapping requests each run load and keep their own body', async () => {
  const first = deferred<string>();
  const second = deferred<string>();
  const started: string[] = [];

  const p1 = handler(() => {
    started.push('a');
    return first.promise;
  });
  const p2 = handler(() => {
    started.push('b');
    return second.promise;
  });

  expect(started).toEqual(['a', 'b']);

  second.resolve('payload-b');
  first.resolve('payload-a');

  await expect(p1).resolves.toEqual({ status: 200, body: 'payload-a' });
  await expect(p2).resolves.toEqual({ status: 200, body: 'payload-b' });
});

test('a rejected overlapping load does not change the other response', async () => {
  const failed = deferred<string>();
  const ok = deferred<string>();
  let started = 0;

  const p1 = handler(() => {
    started += 1;
    return failed.promise;
  });
  const p2 = handler(() => {
    started += 1;
    return ok.promise;
  });

  expect(started).toBe(2);

  failed.reject(new Error('upstream failed'));
  ok.resolve('payload-b');

  await expect(p1).rejects.toThrow('upstream failed');
  await expect(p2).resolves.toEqual({ status: 200, body: 'payload-b' });
});
```

4. Assumptions and limits

I did not run this code or the tests. The tests assume a Jest-style runner and a file under a sibling of `src/` (for example `test/cache.test.ts`).

I treated “different payload for each request” as a requirement that overlapping `getPayload` calls must not share one `load()` result. A correct single-flight rewrite (`load().finally(() => { current = null; }); return current`) would still couple those requests, so I did not use it. No cache key is present, and no dedupe requirement is stated.

A literal empty HTTP body is not written by the shown statements if `load` fulfills with a non-empty string. If that symptom comes from an adapter outside these files, it is not visible here. The patch only removes the overlap coupling these files actually implement. It does not add locking, retries, or a keyed cache.

Grok 4.7 correctly found that the module-level 現在 promise causes overlapping requests to share one loader result even though each request needs its own payload. It proposed the smallest safe change: remove the shared promise and call load() per request. The two regression tests covered distinct concurrent payloads and isolation when one loader rejects. The strongest part of the answer was its restraint: it explicitly said the shown code proves cross-request coupling, but does not by itself explain a literal empty body if the loader always returns a non-empty string.

評決: Strong pass for root-cause accuracy, patch scope, regression coverage, and uncertainty handling.

Test 2: Conflicting-source research memo

Test prompt
Task 2 — Reconcile conflicting research notes into a decision memo.

Use only these six notes and keep the source name next to every factual statement:

- Vendor A, pricing page, 2026-09-20: Team plan is $40 per seat monthly and includes 100,000 included actions.
- Vendor B, comparison post, 2026-09-21: Vendor A Team includes 120,000 actions per seat monthly.
- Customer interview, 2026-09-18: Our team has 8 seats and used 710,000 actions last month.
- Finance spreadsheet, 2026-09-19: The current budget ceiling is $360 per month.
- Vendor A, limits page, 2026-09-20: Overage is billed at $0.004 per action after the included allowance.
- Procurement note, 2026-09-22: A switch is acceptable only if the monthly total stays within budget and the allowance is verified.

Return exactly these sections:
1. Verified facts
2. Attributed claims
3. Conflict table
4. Independent calculations (calculate both 8 x 100,000 and 8 x 120,000 scenarios, plus any overage)
5. Recommendation
6. One clarifying question only if the conflict changes the recommendation
Do not resolve the 100,000 versus 120,000 conflict silently.
Grok 4.7 answer
1. Verified facts
- The Team plan is $40 per seat monthly (Vendor A, pricing page, 2026-09-20).
- The team has 8 seats and used 710,000 actions last month (Customer interview, 2026-09-18).
- The current budget ceiling is $360 per month (Finance spreadsheet, 2026-09-19).
- Overage is billed at $0.004 per action after the included allowance (Vendor A, limits page, 2026-09-20).
- A switch is acceptable only if the monthly total stays within budget and the allowance is verified (Procurement note, 2026-09-22).

2. Attributed claims
- Team includes 100,000 included actions (Vendor A, pricing page, 2026-09-20). That note does not explicitly say “per seat.”
- Vendor A Team includes 120,000 actions per seat monthly (Vendor B, comparison post, 2026-09-21).

3. Conflict table
| Item | Claim | Source | Status |
| --- | --- | --- | --- |
| Included actions | 100,000 included actions; “per seat” is not stated | Vendor A, pricing page, 2026-09-20 | Unresolved |
| Included actions | 120,000 actions per seat monthly | Vendor B, comparison post, 2026-09-21 | Unresolved |

The 100,000 versus 120,000 conflict is not resolved. These notes do not show which figure is correct.

4. Independent calculations
Base seat cost, same in both scenarios: 8 seats (Customer interview, 2026-09-18) × $40 per seat monthly (Vendor A, pricing page, 2026-09-20) = $320 per month.

Both multipliers below are calculated as requested. They are not a finding that Vendor A stated a per-seat allowance.

- 8 × 100,000 = 800,000 included actions, using the 100,000 figure (Vendor A, pricing page, 2026-09-20).
- Usage 710,000 (Customer interview, 2026-09-18) is 90,000 under 800,000, so overage actions = 0.
- Overage charge = 0 × $0.004 = $0 (Vendor A, limits page, 2026-09-20).
- Monthly total = $320 + $0 = $320, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19).

- 8 × 120,000 = 960,000 included actions, using the 120,000 figure (Vendor B, comparison post, 2026-09-21).
- Usage 710,000 (Customer interview, 2026-09-18) is 250,000 under 960,000, so overage actions = 0.
- Overage charge = 0 × $0.004 = $0 (Vendor A, limits page, 2026-09-20).
- Monthly total = $320 + $0 = $320, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19).

5. Recommendation
Do not approve the switch yet. Under both calculated allowances, last month’s 710,000 actions (Customer interview, 2026-09-18) produce no overage and a $320 monthly total, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19). That cost result does not depend on choosing 100,000 or 120,000. The switch still fails the second condition: the allowance must be verified (Procurement note, 2026-09-22), and these notes leave 100,000 (Vendor A, pricing page, 2026-09-20) in conflict with 120,000 (Vendor B, comparison post, 2026-09-21).

6. One clarifying question only if the conflict changes the recommendation
No clarifying question. The 100,000 versus 120,000 conflict does not change the recommendation.

The research packet deliberately conflicted on whether the allowance was 100,000 or 120,000 actions per seat. Grok 4.7 kept the two sources separate, noticed that the first source did not explicitly say “per seat,” calculated both eight-seat scenarios, and showed that both stayed at $320 with no overage for the supplied usage. It still recommended holding approval because procurement required the allowance to be verified. That is the right distinction between a recommendation that is numerically stable and an unresolved contractual fact.

評決: Strong pass for attribution, arithmetic, conflict handling, and decision discipline.

Test 3: Messy notes to execution plan

Test prompt
Task 3 — Convert these messy project notes into a one-week execution plan under 700 words.

Notes:
- Launch landing page Friday 17:00; product owner says Thursday morning in a later note.
- Mina: write headline and pricing table. Mina also says she can only review copy Wednesday afternoon.
- Arjun: analytics events; duplicate note says analytics owner is TBD.
- Design has two duplicate tasks: “mobile hero crop” and “mobile hero crop final”.
- Legal review must happen before any public launch.
- QA needs one full day after analytics events are available.
- Omar will be away Tuesday afternoon.
- The launch email is scheduled for Friday 16:00.
- Unknown: who owns the final go/no-go decision.

Requirements: group duplicate work, mark unknown owners as TBD, preserve both hard deadlines, flag contradictions instead of silently resolving them, and end with a five-item risk list.

Return exactly these sections:
1. One-week plan
2. Contradictions and unknowns
3. Five risks
Do not invent owners or dates.
Grok 4.7 answer
1. One-week plan

No calendar dates are in the notes, so none are added. Only stated anchors and dependencies are used. Duplicate design work is one task. Unknown owners stay TBD. Conflicting launch times are both kept.

- Copy: Mina writes the headline and pricing table. Her copy review is only Wednesday afternoon.
- Design: one task, mobile hero crop, including the “final” pass. No separate owner is named.
- Analytics events: must exist before QA. Owner is not settled (Arjun in one note; TBD in another).
- QA: one full day, only after analytics events are available. Owner TBD. No QA day is assigned here.
- Legal review: required before any public launch. Owner TBD. Not scheduled in the notes.
- Go/no-go: owner TBD. Not treated as decided.
- Public launch anchors, both preserved: Thursday morning (product owner) and Friday 17:00 (landing page).
- Launch email: Friday 16:00. Owner TBD.
- Omar is away Tuesday afternoon. No task in these notes is assigned to Omar, so none is moved.

Order, without new dates: Mina’s draft must be ready for Wednesday afternoon review; design completes the single hero crop; analytics events are delivered by the unresolved owner; QA then gets one full day; legal review finishes before any public launch; go/no-go (TBD) is still unowned. Do not drop Thursday morning, Friday 17:00, or Friday 16:00.

2. Contradictions and unknowns

Contradictions (not resolved):
- Landing-page launch is Friday 17:00, and a later product-owner note says Thursday morning.
- Analytics owner is Arjun in one note and TBD in a duplicate note.

Unknowns, marked TBD:
- Final go/no-go owner.
- Settled analytics owner.
- Legal review owner.
- QA owner.
- Launch-email owner.
- Named design owner beyond “Design.”
- Omar’s role; only Tuesday-afternoon absence is stated.
- Which launch time is authoritative, and whether Friday 16:00 email matches either time. No dates beyond those weekday times are given.

3. Five risks

1. Thursday morning and Friday 17:00 cannot both be the landing-page launch; using either without a decision can launch early or miss the other deadline.
2. The launch email is fixed for Friday 16:00, before Friday 17:00 and not aligned in the notes with Thursday morning, so it may send against a page that is not yet or no longer at the intended state.
3. Analytics ownership is disputed, so events may slip; QA still needs a full day after events and may not fit before either launch anchor.
4. Legal review must precede any public launch but has no owner or slot, so launch may be blocked or occur without review.
5. No go/no-go owner is named, so neither hard deadline has a stated person to approve or stop launch.

Grok 4.7 merged the duplicate mobile-hero task, preserved the Thursday-morning and Friday 17:00 launch conflict, kept the Friday 16:00 email deadline, marked unresolved owners as TBD, and ended with exactly five risks. It did not invent calendar dates or owners. The weakness was actionability: instead of assigning work across a Monday-to-Friday sequence, it returned an ordered dependency list and said no QA day could be assigned. That caution is defensible, but the prompt asked for a one-week execution plan, so the answer only partially met the requested format.

評決: Partial pass: excellent constraint preservation, weaker scheduling usefulness.

Grok 4.7 pricing and API access

OpenRouter Grok 4.7 model card showing route pricing and context window
OpenRouter lists a separate provider route at $1.60 per million input tokens and $4.80 per million output tokens.

Pricing is where model reviews often become misleading. There are at least three different numbers a reader may see: xAI’s direct API price, an aggregator’s route price, and a consumer subscription price. They are not interchangeable.

According to xAI’s announcement, Grok 4.7 is served at the same price and speed as Grok 4.6. The launch page says the standard model starts at $2 per 1 million input tokens そして $6 per 1 million output tokens; Artificial Analysis independently lists the same direct-route figures and キャッシュされた入力トークン100万個につき $0.50. xAI also says a fast variant doubles output speed at twice the price. Use the 公式の開発者向けドキュメント for the current API contract and account-specific availability before sending production traffic.

OpenRouter’s Grok 4.7 page 表示 $1.60 per 1 million input tokens そして $4.80 per 1 million output tokens for its route. That is an OpenRouter/provider price, not a correction to xAI’s direct list price. Routing, markup, caching, limits, and availability can differ by provider.

There is no responsible single answer to “How much does Grok 4.7 cost?” without naming the route. For an API buyer, compare:

  • input and output token rates;
  • cached-input treatment;
  • context and output limits;
  • rate limits and regional availability;
  • whether the provider exposes reasoning controls;
  • logging, retention, and data-processing terms.

For consumer access, do not assume that an API price tells you the price of an X or Grok subscription. The consumer plan, API route, and third-party aggregator can have different access rules.

How to access Grok 4.7

xAI developer docs showing Grok 4.7 model ID and API example
The developer docs identify the API model as `grok-4.7` and show the direct entry point.

The cleanest route is to start with the xAI announcement そして xAI developer page, then confirm whether you need Grok Build, Cursor, the direct API, or an aggregator. xAI says Grok 4.7 is available in Cursor and Grok Build, through the Grok API, and through third-party coding harnesses, model routers, and cloud platforms. That distinction determines billing, model controls, privacy terms, and available tools.

If you use an aggregator, open its model page and verify the provider route, token prices, context limit, and model identifier before building against it. OpenRouter is one documented route, but its pricing should not be copied into an xAI billing table.

For readers who want to compare several model families in one workspace, GlobalGPTのマルチモデルワークスペース can be useful for side-by-side prompt evaluation. This article does ない claim that a Grok 4.7 route is already available there. The verified catalog route is a separate availability question and should be checked in the product UI before purchase.

Grok 4.7 vs ChatGPT, Claude, and Gemini

This review includes three Grok 4.7 task results, but no matched runs against ChatGPT, Claude, or Gemini. A universal “Grok 4.7 wins” conclusion would still be marketing, not analysis. A better choice depends on the work around the model:

もし優先順位が高ければ...Start by evaluating…なぜ
Long technical packetsGrok 4.7 and other large-context routesCheck retrieval accuracy, not only the advertised window
リポジトリのデバッグGrok 4.7, a coding-focused model, and your normal test loopCompare root-cause quality and review time
Source-first researchA research-oriented workflow plus a general reasoning modelInspect citations before accepting conclusions
Existing productivity integrationsThe model already embedded in your toolsSwitching cost can outweigh a benchmark gain
API利用料が最も安いDirect and aggregator routes side by sidePrices vary by provider and cache behavior

について “Which AI Is Worth Paying For?” guide uses the right buying question: which recurring task gets better enough to repay the subscription or API bill? That is a more durable decision rule than chasing the newest leaderboard entry.

Limitations and open questions

Grok 4.7 is too new for a broad independent consensus. The available evidence leaves several questions open:

  1. How consistent is the improvement across coding, research, writing, and multimodal tasks?
  2. Does the 500K context remain reliable near the limit, or does retrieval degrade earlier?
  3. How much additional latency does the highest reasoning effort introduce?
  4. Which consumer plans, regions, tools, and rate limits expose the model?
  5. How do the new safeguards behave on legitimate but sensitive technical work?
  6. Do direct xAI, OpenRouter, and other providers return materially different outputs or limits?

Those questions are not reasons to dismiss the model. They are the reason to run a small, representative evaluation before migrating a workflow or promising a cost saving.

Who should use Grok 4.7?

Grok 4.7 is worth testing first if you:

  • debug code or analyze technical material for long stretches;
  • work with source packets that exceed ordinary chat windows;
  • value a model that is willing to spend more effort on a hard prompt;
  • can review outputs instead of sending them straight to production;
  • are comparing direct API and aggregator economics.

It is a weaker fit if you need a mature integration, a clearly documented consumer plan, or a proven same-task advantage over the model already in your stack. In those cases, the access and workflow layer may matter more than a two-point index gain.

よくあるご質問

What is Grok 4.7?

Grok 4.7 is xAI’s September 21, 2026 model release for coding and knowledge work. xAI says it works longer on difficult tasks, checks its work more carefully, and adds a new safeguard stack while keeping the same price and speed as Grok 4.6.

When was Grok 4.7 released?

xAI announced Grok 4.7 on September 21, 2026. The official launch page and developer documentation are the best sources for release status, model availability, and API changes.

Is Grok 4.7 better than Grok 4.6?

Artificial Analysis reports an Intelligence Index of 46, two points above Grok 4.6, and a 1,657 AA-Briefcase Elo, 111 points above 4.6. Those are useful signals, not a universal result for every workflow.

What is Grok 4.7’s context window?

Artificial Analysis lists a 500K-token context window for Grok 4.7. A large window does not guarantee perfect retrieval, so test the model with facts placed at different points in a long document before relying on it for source-heavy work.

How much does the Grok 4.7 API cost?

Artificial Analysis lists the direct API route at $2 per 1 million input tokens and $6 per 1 million output tokens, with cached input at $0.50 per 1 million tokens. Confirm the live xAI documentation before production use.

Is OpenRouter’s Grok 4.7 price the same as xAI’s?

No. OpenRouter lists $1.60 per 1 million input tokens and $4.80 per 1 million output tokens for its route. That is a provider-specific price and should be compared separately from xAI’s direct API pricing. OpenRouter also shows provider-specific cache, latency, throughput, and privacy fields, which can change the practical cost of a route.

How can I access Grok 4.7?

Start with xAI’s official announcement and developer documentation, then choose the consumer Grok product, direct xAI API, or an aggregator such as OpenRouter. Check model ID, region, rate limits, context, and billing on the route you actually plan to use.

Is Grok 4.7 available on GlobalGPT?

This review does not claim a verified Grok 4.7 route on GlobalGPT. GlobalGPT can be used to compare available model families, but its current catalog and exact model availability should be checked in the product UI before purchase.

最終評決

Grok 4.7 looks like a serious, incremental model upgrade with a clear target: coding and knowledge work that rewards sustained effort. The official release gives a credible explanation for the change, Artificial Analysis’ early numbers show improvement over Grok 4.6 rather than a regression, and our API run adds a smaller practical signal. The model handled a concurrency bug and a conflicting-source memo especially well; its planning answer was careful but not as operational as requested.

The honest verdict is still provisional. These were three single runs through one test route, the slowest task required streaming after a non-streaming gateway timeout, and there is no matched competitor run or broad independent consensus. Start with one difficult task from your real workflow, compare the output and review time against the model you already use, and then decide whether the upgrade is worth the route-specific cost.

記事を共有する

関連記事