Grok 4.7 is a serious coding and knowledge-work upgrade, but the evidence is still early: official launch claims and independent benchmarks point to improvement, not a universal win.
Réponse rapide
Grok 4.7 is xAI’s September 21, 2026 upgrade aimed at coding and difficult knowledge-work tasks. xAI says it spends longer on hard problems, checks its work more carefully, and adds a new safeguard stack while keeping the same price and speed as Grok 4.6. The launch page also presents it as twice as fast and half the price of comparable models, a vendor comparison that needs to be read alongside the underlying task and route. Independent Artificial Analysis data gives it a small overall lift over Grok 4.6 and records a 500K context window. In our three-task API test, Grok 4.7 was strongest at debugging and source-boundary discipline; its project plan preserved uncertainty well but was less actionable than requested. That is useful hands-on evidence, not proof that it beats every competing model.

What is Grok 4.7?

Grok 4.7 is the latest model in xAI’s Grok line, announced on September 21, 2026. The official xAI announcement positions it as the company’s most capable model for coding and knowledge work. The meaningful change is not a new consumer-facing feature list. It is a change in how the model is expected to spend effort: more time on difficult tasks, more self-checking, and a safety layer intended to make that extra capability more controllable.
That positioning matters because it describes a workflow model rather than a chatbot personality. Grok 4.7 is being sold as a model for tasks that benefit from sustained analysis: debugging, planning, technical writing, research synthesis, and long prompts that need the model to keep several constraints in view.
The boundary is equally important. xAI’s announcement is a vendor statement about intended behavior. It is not an independent proof that Grok 4.7 beats every competing model, every prompt, or every coding repository. Treat the launch page as the source for what xAI claims, then use independent measurements and your own task to decide whether those claims matter.
Grok 4.7 vs Grok 4.6: what actually changed?

The most defensible summary is an effort and reliability upgrade, not a completely different product.
| Zone | Grok 4.7 | What the evidence supports |
|---|---|---|
| Sortie | September 21, 2026 | Date stated on the official xAI launch page |
| Main positioning | Coding and knowledge work | Official xAI positioning |
| Difficult tasks | Longer effort and more careful self-checking | Official xAI description; our three-task API run offers a small hands-on check, not a matched comparison |
| Sécurité | New safeguard stack | Official xAI description |
| Price and speed | Same as Grok 4.6, according to xAI | A launch claim, not a third-party price audit |
| Contexte | 500K tokens in the Artificial Analysis profile | Independent model profile; verify the route you use |
| Reasoning controls | Low through xhigh in Artificial Analysis’ listing | Independent route/model metadata |
This is a useful distinction for buyers. If Grok 4.6 already fit your workflow, 4.7 is a low-friction upgrade to evaluate. If 4.6 did not fit because of ecosystem, latency, reliability, or access constraints, a higher benchmark number alone does not solve those problems.
For a broader model-buying framework, the Guide des tarifs GPT-5.6 et Comparaison GPT-5.6 are useful reminders to separate model quality from plan economics and workflow fit.
Capabilities: where Grok 4.7 should help
Codage et débogage
xAI’s strongest launch framing is coding. The promise is not simply that Grok 4.7 writes more code; it is that it can stay with a hard engineering problem longer, inspect its own answer, and return a more considered patch or explanation. That is the behavior to look for in a real repository: does it preserve the existing design, identify the root cause, add useful tests, and state what it still cannot know?
Do not evaluate the model with a greenfield “build a todo app” prompt. Give it a bounded bug with a failing test, a small set of files, and explicit constraints. Ask for a diagnosis before a patch. This makes it easier to see whether the longer reasoning claim produces less review work or merely longer prose.
The same rule applies to teams comparing different coding routes. A Test du DeepSeek V4 Pro et le Examen du plan de codage GLM provide useful comparison context, but neither should be treated as a substitute for testing the repository and toolchain you actually use.
Long-context knowledge work
Artificial Analysis lists a 500K-token context window for Grok 4.7. That is large enough to change how you approach source-heavy work: a long technical specification, a set of meeting notes, a codebase slice, or a research packet can stay in one conversation instead of being repeatedly summarized.
Context capacity is not the same as context quality. A model can accept a large input and still miss a constraint near the middle, over-weight the latest paragraph, or produce a confident summary that needs source checking. The practical test is retrieval under pressure: place a few decisive facts in different parts of a long document, ask for a constrained decision, and check whether the answer cites the right passages.
SpaceXAI’s official X account repeated the launch framing on September 22, calling Grok 4.7 a notable improvement over Grok 4.6 at the same price and speed and sharing the model comparison graphic. That post is useful evidence of the company’s public positioning, not independent validation or a community consensus.
If the model is used for research, keep the evidence loop visible. Our Comparaison entre Gemini et Perplexity et Guide des alternatives à Perplexity make the same practical point: a polished answer is not a source audit.
Self-checking and safety
xAI says Grok 4.7 checks its work more carefully and introduces a new safeguard stack. Those are valuable goals, especially when the model is asked to take a long chain of decisions. They are not a guarantee that the output is correct or safe for automatic execution.
For production use, keep ordinary controls around the model: tests, permissions, human review, citation checks, and a clear stop condition. The model’s self-check is another signal to inspect, not a replacement for the controls around it.
Grok 4.7 benchmarks: what the numbers prove


The most useful independent snapshot currently comes from Artificial Analysis’ Grok 4.7 benchmarking article, dated September 21, 2026.
| Mesure | Résultat publié | Comment le lire |
|---|---|---|
| Indice d'analyse par intelligence artificielle | 46 | Two points higher than Grok 4.6 in the same index |
| AA-Porte-documents | 1,657 Elo | 111 Elo above Grok 4.6 in that evaluation |
| Fenêtre de contexte | 500 000 jetons | A model-profile field, not proof of perfect long-context retrieval |
| Effort de raisonnement | Low to xhigh | A route-level control listed by Artificial Analysis |
The direction is encouraging: the independent index shows improvement over the previous version, and the larger Elo gap in AA-Briefcase suggests a more noticeable gain on that particular evaluation. But the table still describes two datasets, not a universal ranking. Benchmark composition, tool access, sampling, and model settings can change the result.
The official xAI launch page shows Grok 4.7 at 46.3% on CursorBench 4.0 in its comparison table, alongside Grok 4.6 at 40.4%, GPT-5.6 Sol at 41.7%, and Fable 5.1 at 51.8%. The page also shows 71.0% on DeepSWE v1.1 at high effort, 64.0% on EEBench, and 1,657 on AA-Briefcase v1.1. These are vendor-presented comparison figures, so treat them as official reported results rather than a substitute for independent replication. The safer conclusion is that xAI is targeting serious coding and professional knowledge work, while independent measurements show a meaningful improvement on some tasks and flatter results on others.
Hands-on Grok 4.7 test: three practical tasks
We ran three fixed prompts against the grok-4.7 model ID through an OpenAI-compatible test API on September 23, 2026. This was a single-route evaluation, not the direct xAI API and not a matched competitor test. The visible answer from each run is preserved below; private reasoning traces are excluded.
| Tâche | Heure observée | Résultat |
|---|---|---|
| TypeScript concurrency bug | 333.7 seconds, streamed | Strong pass: correct shared-promise diagnosis, minimal patch, two useful tests, and a careful limit on what the snippet could prove |
| Conflicting research notes | 68.7 seconds | Strong pass: preserved attribution, calculated both allowance scenarios, and refused approval until the source conflict was verified |
| Messy one-week plan | 50.6 seconds | Partial pass: preserved deadlines, contradictions, unknown owners, and five risks, but did not turn the work into a clear day-by-day schedule |
The latency numbers describe this route and these prompts only. The coding task first hit a non-streaming gateway timeout, then completed with streaming enabled. That is a practical deployment warning about transport and timeout handling, not a general Grok 4.7 speed score.
Test 1: TypeScript concurrency debugging
Task 1 — TypeScript API handler concurrency debugging
A small TypeScript handler intermittently returns HTTP 200 with an empty body when two requests arrive close together.
Files provided:
src/cache.ts
```ts
let current: Promise<string> | null = null;
export async function getPayload(load: () => Promise<string>): Promise<string> {
if (!current) {
current = load();
}
const value = await current;
current = null;
return value;
}
```
src/handler.ts
```ts
import { getPayload } from './cache';
export async function handler(load: () => Promise<string>) {
const payload = await getPayload(load);
return { status: 200, body: payload };
}
```
The production symptom is intermittent only when two requests overlap. The loader itself is safe to call concurrently and returns a different payload for each request. No other files are available.
Return exactly these sections:
1. Root cause
2. Smallest safe patch (show a unified diff)
3. Two regression tests (TypeScript or Jest-style code)
4. Assumptions and limits
Do not claim to have run the code.1. Root cause
`current` is module-level state shared by every call. On one JavaScript turn the null check and assignment are atomic, so two overlapping calls do not both run `load()`. The first call stores its promise; a second call that arrives before that promise settles skips its own `load()` and awaits the same promise. `handler` then returns HTTP 200 whose `body` is the first call’s string, not the second request’s payload. That window exists only while a call is in flight, so the failure shows up only when requests overlap.
`current = null` runs in each continuation after `await` has already captured the promise. It does not make the other waiter re-read the variable, so these files alone do not resolve that waiter to `null` or `""`. What they do produce on overlap is a shared resolution and a skipped per-request load. The clear also makes the slot a bad single-flight cache: the first continuation to resume drops `current`, and a later caller can start a different `load()` while another caller is still unwinding.
2. Smallest safe patch (show a unified diff)
Each request must call `load()` itself. Concurrent `load()` calls are allowed, and each returns a different payload, so deduping on one shared promise is unsafe. `handler.ts` does not need a change.
```diff
--- a/src/cache.ts
+++ b/src/cache.ts
@@ -1,10 +1,3 @@
-let current: Promise<string> | null = null;
-
export async function getPayload(load: () => Promise<string>): Promise<string> {
- if (!current) {
- current = load();
- }
- const value = await current;
- current = null;
- return value;
+ return await load();
}
```
3. Two regression tests (TypeScript or Jest-style code)
```ts
import { handler } from '../src/handler';
function deferred<T>() {
let resolve!: (value: T) => void;
let reject!: (reason: unknown) => void;
const promise = new Promise<T>((res, rej) => {
resolve = res;
reject = rej;
});
return { promise, resolve, reject };
}
test('overlapping requests each run load and keep their own body', async () => {
const first = deferred<string>();
const second = deferred<string>();
const started: string[] = [];
const p1 = handler(() => {
started.push('a');
return first.promise;
});
const p2 = handler(() => {
started.push('b');
return second.promise;
});
expect(started).toEqual(['a', 'b']);
second.resolve('payload-b');
first.resolve('payload-a');
await expect(p1).resolves.toEqual({ status: 200, body: 'payload-a' });
await expect(p2).resolves.toEqual({ status: 200, body: 'payload-b' });
});
test('a rejected overlapping load does not change the other response', async () => {
const failed = deferred<string>();
const ok = deferred<string>();
let started = 0;
const p1 = handler(() => {
started += 1;
return failed.promise;
});
const p2 = handler(() => {
started += 1;
return ok.promise;
});
expect(started).toBe(2);
failed.reject(new Error('upstream failed'));
ok.resolve('payload-b');
await expect(p1).rejects.toThrow('upstream failed');
await expect(p2).resolves.toEqual({ status: 200, body: 'payload-b' });
});
```
4. Assumptions and limits
I did not run this code or the tests. The tests assume a Jest-style runner and a file under a sibling of `src/` (for example `test/cache.test.ts`).
I treated “different payload for each request” as a requirement that overlapping `getPayload` calls must not share one `load()` result. A correct single-flight rewrite (`load().finally(() => { current = null; }); return current`) would still couple those requests, so I did not use it. No cache key is present, and no dedupe requirement is stated.
A literal empty HTTP body is not written by the shown statements if `load` fulfills with a non-empty string. If that symptom comes from an adapter outside these files, it is not visible here. The patch only removes the overlap coupling these files actually implement. It does not add locking, retries, or a keyed cache.Grok 4.7 correctly found that the module-level actuel promise causes overlapping requests to share one loader result even though each request needs its own payload. It proposed the smallest safe change: remove the shared promise and call load() per request. The two regression tests covered distinct concurrent payloads and isolation when one loader rejects. The strongest part of the answer was its restraint: it explicitly said the shown code proves cross-request coupling, but does not by itself explain a literal empty body if the loader always returns a non-empty string.
Verdict : Strong pass for root-cause accuracy, patch scope, regression coverage, and uncertainty handling.
Test 2: Conflicting-source research memo
Task 2 — Reconcile conflicting research notes into a decision memo. Use only these six notes and keep the source name next to every factual statement: - Vendor A, pricing page, 2026-09-20: Team plan is $40 per seat monthly and includes 100,000 included actions. - Vendor B, comparison post, 2026-09-21: Vendor A Team includes 120,000 actions per seat monthly. - Customer interview, 2026-09-18: Our team has 8 seats and used 710,000 actions last month. - Finance spreadsheet, 2026-09-19: The current budget ceiling is $360 per month. - Vendor A, limits page, 2026-09-20: Overage is billed at $0.004 per action after the included allowance. - Procurement note, 2026-09-22: A switch is acceptable only if the monthly total stays within budget and the allowance is verified. Return exactly these sections: 1. Verified facts 2. Attributed claims 3. Conflict table 4. Independent calculations (calculate both 8 x 100,000 and 8 x 120,000 scenarios, plus any overage) 5. Recommendation 6. One clarifying question only if the conflict changes the recommendation Do not resolve the 100,000 versus 120,000 conflict silently.
1. Verified facts - The Team plan is $40 per seat monthly (Vendor A, pricing page, 2026-09-20). - The team has 8 seats and used 710,000 actions last month (Customer interview, 2026-09-18). - The current budget ceiling is $360 per month (Finance spreadsheet, 2026-09-19). - Overage is billed at $0.004 per action after the included allowance (Vendor A, limits page, 2026-09-20). - A switch is acceptable only if the monthly total stays within budget and the allowance is verified (Procurement note, 2026-09-22). 2. Attributed claims - Team includes 100,000 included actions (Vendor A, pricing page, 2026-09-20). That note does not explicitly say “per seat.” - Vendor A Team includes 120,000 actions per seat monthly (Vendor B, comparison post, 2026-09-21). 3. Conflict table | Item | Claim | Source | Status | | --- | --- | --- | --- | | Included actions | 100,000 included actions; “per seat” is not stated | Vendor A, pricing page, 2026-09-20 | Unresolved | | Included actions | 120,000 actions per seat monthly | Vendor B, comparison post, 2026-09-21 | Unresolved | The 100,000 versus 120,000 conflict is not resolved. These notes do not show which figure is correct. 4. Independent calculations Base seat cost, same in both scenarios: 8 seats (Customer interview, 2026-09-18) × $40 per seat monthly (Vendor A, pricing page, 2026-09-20) = $320 per month. Both multipliers below are calculated as requested. They are not a finding that Vendor A stated a per-seat allowance. - 8 × 100,000 = 800,000 included actions, using the 100,000 figure (Vendor A, pricing page, 2026-09-20). - Usage 710,000 (Customer interview, 2026-09-18) is 90,000 under 800,000, so overage actions = 0. - Overage charge = 0 × $0.004 = $0 (Vendor A, limits page, 2026-09-20). - Monthly total = $320 + $0 = $320, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19). - 8 × 120,000 = 960,000 included actions, using the 120,000 figure (Vendor B, comparison post, 2026-09-21). - Usage 710,000 (Customer interview, 2026-09-18) is 250,000 under 960,000, so overage actions = 0. - Overage charge = 0 × $0.004 = $0 (Vendor A, limits page, 2026-09-20). - Monthly total = $320 + $0 = $320, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19). 5. Recommendation Do not approve the switch yet. Under both calculated allowances, last month’s 710,000 actions (Customer interview, 2026-09-18) produce no overage and a $320 monthly total, which is within the $360 ceiling (Finance spreadsheet, 2026-09-19). That cost result does not depend on choosing 100,000 or 120,000. The switch still fails the second condition: the allowance must be verified (Procurement note, 2026-09-22), and these notes leave 100,000 (Vendor A, pricing page, 2026-09-20) in conflict with 120,000 (Vendor B, comparison post, 2026-09-21). 6. One clarifying question only if the conflict changes the recommendation No clarifying question. The 100,000 versus 120,000 conflict does not change the recommendation.
The research packet deliberately conflicted on whether the allowance was 100,000 or 120,000 actions per seat. Grok 4.7 kept the two sources separate, noticed that the first source did not explicitly say “per seat,” calculated both eight-seat scenarios, and showed that both stayed at $320 with no overage for the supplied usage. It still recommended holding approval because procurement required the allowance to be verified. That is the right distinction between a recommendation that is numerically stable and an unresolved contractual fact.
Verdict : Strong pass for attribution, arithmetic, conflict handling, and decision discipline.
Test 3: Messy notes to execution plan
Task 3 — Convert these messy project notes into a one-week execution plan under 700 words. Notes: - Launch landing page Friday 17:00; product owner says Thursday morning in a later note. - Mina: write headline and pricing table. Mina also says she can only review copy Wednesday afternoon. - Arjun: analytics events; duplicate note says analytics owner is TBD. - Design has two duplicate tasks: “mobile hero crop” and “mobile hero crop final”. - Legal review must happen before any public launch. - QA needs one full day after analytics events are available. - Omar will be away Tuesday afternoon. - The launch email is scheduled for Friday 16:00. - Unknown: who owns the final go/no-go decision. Requirements: group duplicate work, mark unknown owners as TBD, preserve both hard deadlines, flag contradictions instead of silently resolving them, and end with a five-item risk list. Return exactly these sections: 1. One-week plan 2. Contradictions and unknowns 3. Five risks Do not invent owners or dates.
1. One-week plan No calendar dates are in the notes, so none are added. Only stated anchors and dependencies are used. Duplicate design work is one task. Unknown owners stay TBD. Conflicting launch times are both kept. - Copy: Mina writes the headline and pricing table. Her copy review is only Wednesday afternoon. - Design: one task, mobile hero crop, including the “final” pass. No separate owner is named. - Analytics events: must exist before QA. Owner is not settled (Arjun in one note; TBD in another). - QA: one full day, only after analytics events are available. Owner TBD. No QA day is assigned here. - Legal review: required before any public launch. Owner TBD. Not scheduled in the notes. - Go/no-go: owner TBD. Not treated as decided. - Public launch anchors, both preserved: Thursday morning (product owner) and Friday 17:00 (landing page). - Launch email: Friday 16:00. Owner TBD. - Omar is away Tuesday afternoon. No task in these notes is assigned to Omar, so none is moved. Order, without new dates: Mina’s draft must be ready for Wednesday afternoon review; design completes the single hero crop; analytics events are delivered by the unresolved owner; QA then gets one full day; legal review finishes before any public launch; go/no-go (TBD) is still unowned. Do not drop Thursday morning, Friday 17:00, or Friday 16:00. 2. Contradictions and unknowns Contradictions (not resolved): - Landing-page launch is Friday 17:00, and a later product-owner note says Thursday morning. - Analytics owner is Arjun in one note and TBD in a duplicate note. Unknowns, marked TBD: - Final go/no-go owner. - Settled analytics owner. - Legal review owner. - QA owner. - Launch-email owner. - Named design owner beyond “Design.” - Omar’s role; only Tuesday-afternoon absence is stated. - Which launch time is authoritative, and whether Friday 16:00 email matches either time. No dates beyond those weekday times are given. 3. Five risks 1. Thursday morning and Friday 17:00 cannot both be the landing-page launch; using either without a decision can launch early or miss the other deadline. 2. The launch email is fixed for Friday 16:00, before Friday 17:00 and not aligned in the notes with Thursday morning, so it may send against a page that is not yet or no longer at the intended state. 3. Analytics ownership is disputed, so events may slip; QA still needs a full day after events and may not fit before either launch anchor. 4. Legal review must precede any public launch but has no owner or slot, so launch may be blocked or occur without review. 5. No go/no-go owner is named, so neither hard deadline has a stated person to approve or stop launch.
Grok 4.7 merged the duplicate mobile-hero task, preserved the Thursday-morning and Friday 17:00 launch conflict, kept the Friday 16:00 email deadline, marked unresolved owners as TBD, and ended with exactly five risks. It did not invent calendar dates or owners. The weakness was actionability: instead of assigning work across a Monday-to-Friday sequence, it returned an ordered dependency list and said no QA day could be assigned. That caution is defensible, but the prompt asked for a one-week execution plan, so the answer only partially met the requested format.
Verdict : Partial pass: excellent constraint preservation, weaker scheduling usefulness.
Grok 4.7 pricing and API access

Pricing is where model reviews often become misleading. There are at least three different numbers a reader may see: xAI’s direct API price, an aggregator’s route price, and a consumer subscription price. They are not interchangeable.
According to xAI’s announcement, Grok 4.7 is served at the same price and speed as Grok 4.6. The launch page says the standard model starts at $2 per 1 million input tokens et $6 per 1 million output tokens; Artificial Analysis independently lists the same direct-route figures and $0,50 par million de jetons d'entrée mis en cache. xAI also says a fast variant doubles output speed at twice the price. Use the documentation officielle destinée aux développeurs for the current API contract and account-specific availability before sending production traffic.
OpenRouter’s Grok 4.7 page présente $1.60 per 1 million input tokens et $4.80 per 1 million output tokens for its route. That is an OpenRouter/provider price, not a correction to xAI’s direct list price. Routing, markup, caching, limits, and availability can differ by provider.
There is no responsible single answer to “How much does Grok 4.7 cost?” without naming the route. For an API buyer, compare:
- input and output token rates;
- cached-input treatment;
- context and output limits;
- rate limits and regional availability;
- whether the provider exposes reasoning controls;
- logging, retention, and data-processing terms.
For consumer access, do not assume that an API price tells you the price of an X or Grok subscription. The consumer plan, API route, and third-party aggregator can have different access rules.
How to access Grok 4.7

The cleanest route is to start with the xAI announcement et xAI developer page, then confirm whether you need Grok Build, Cursor, the direct API, or an aggregator. xAI says Grok 4.7 is available in Cursor and Grok Build, through the Grok API, and through third-party coding harnesses, model routers, and cloud platforms. That distinction determines billing, model controls, privacy terms, and available tools.
If you use an aggregator, open its model page and verify the provider route, token prices, context limit, and model identifier before building against it. OpenRouter is one documented route, but its pricing should not be copied into an xAI billing table.
For readers who want to compare several model families in one workspace, L'espace de travail multimodèle de GlobalGPT can be useful for side-by-side prompt evaluation. This article does pas claim that a Grok 4.7 route is already available there. The verified catalog route is a separate availability question and should be checked in the product UI before purchase.
Grok 4.7 vs ChatGPT, Claude, and Gemini
This review includes three Grok 4.7 task results, but no matched runs against ChatGPT, Claude, or Gemini. A universal “Grok 4.7 wins” conclusion would still be marketing, not analysis. A better choice depends on the work around the model:
| Si votre priorité est... | Start by evaluating… | Pourquoi |
|---|---|---|
| Long technical packets | Grok 4.7 and other large-context routes | Check retrieval accuracy, not only the advertised window |
| Débogage du référentiel | Grok 4.7, a coding-focused model, and your normal test loop | Compare root-cause quality and review time |
| Source-first research | A research-oriented workflow plus a general reasoning model | Inspect citations before accepting conclusions |
| Existing productivity integrations | The model already embedded in your tools | Switching cost can outweigh a benchmark gain |
| Facture API la plus basse | Direct and aggregator routes side by side | Prices vary by provider and cache behavior |
Le “Which AI Is Worth Paying For?” guide uses the right buying question: which recurring task gets better enough to repay the subscription or API bill? That is a more durable decision rule than chasing the newest leaderboard entry.
Limitations and open questions
Grok 4.7 is too new for a broad independent consensus. The available evidence leaves several questions open:
- How consistent is the improvement across coding, research, writing, and multimodal tasks?
- Does the 500K context remain reliable near the limit, or does retrieval degrade earlier?
- How much additional latency does the highest reasoning effort introduce?
- Which consumer plans, regions, tools, and rate limits expose the model?
- How do the new safeguards behave on legitimate but sensitive technical work?
- Do direct xAI, OpenRouter, and other providers return materially different outputs or limits?
Those questions are not reasons to dismiss the model. They are the reason to run a small, representative evaluation before migrating a workflow or promising a cost saving.
Who should use Grok 4.7?
Grok 4.7 is worth testing first if you:
- debug code or analyze technical material for long stretches;
- work with source packets that exceed ordinary chat windows;
- value a model that is willing to spend more effort on a hard prompt;
- can review outputs instead of sending them straight to production;
- are comparing direct API and aggregator economics.
It is a weaker fit if you need a mature integration, a clearly documented consumer plan, or a proven same-task advantage over the model already in your stack. In those cases, the access and workflow layer may matter more than a two-point index gain.
FAQ
What is Grok 4.7?
Grok 4.7 is xAI’s September 21, 2026 model release for coding and knowledge work. xAI says it works longer on difficult tasks, checks its work more carefully, and adds a new safeguard stack while keeping the same price and speed as Grok 4.6.
When was Grok 4.7 released?
xAI announced Grok 4.7 on September 21, 2026. The official launch page and developer documentation are the best sources for release status, model availability, and API changes.
Is Grok 4.7 better than Grok 4.6?
Artificial Analysis reports an Intelligence Index of 46, two points above Grok 4.6, and a 1,657 AA-Briefcase Elo, 111 points above 4.6. Those are useful signals, not a universal result for every workflow.
What is Grok 4.7’s context window?
Artificial Analysis lists a 500K-token context window for Grok 4.7. A large window does not guarantee perfect retrieval, so test the model with facts placed at different points in a long document before relying on it for source-heavy work.
How much does the Grok 4.7 API cost?
Artificial Analysis lists the direct API route at $2 per 1 million input tokens and $6 per 1 million output tokens, with cached input at $0.50 per 1 million tokens. Confirm the live xAI documentation before production use.
Is OpenRouter’s Grok 4.7 price the same as xAI’s?
No. OpenRouter lists $1.60 per 1 million input tokens and $4.80 per 1 million output tokens for its route. That is a provider-specific price and should be compared separately from xAI’s direct API pricing. OpenRouter also shows provider-specific cache, latency, throughput, and privacy fields, which can change the practical cost of a route.
How can I access Grok 4.7?
Start with xAI’s official announcement and developer documentation, then choose the consumer Grok product, direct xAI API, or an aggregator such as OpenRouter. Check model ID, region, rate limits, context, and billing on the route you actually plan to use.
Is Grok 4.7 available on GlobalGPT?
This review does not claim a verified Grok 4.7 route on GlobalGPT. GlobalGPT can be used to compare available model families, but its current catalog and exact model availability should be checked in the product UI before purchase.
Verdict final
Grok 4.7 looks like a serious, incremental model upgrade with a clear target: coding and knowledge work that rewards sustained effort. The official release gives a credible explanation for the change, Artificial Analysis’ early numbers show improvement over Grok 4.6 rather than a regression, and our API run adds a smaller practical signal. The model handled a concurrency bug and a conflicting-source memo especially well; its planning answer was careful but not as operational as requested.
The honest verdict is still provisional. These were three single runs through one test route, the slowest task required streaming after a non-streaming gateway timeout, and there is no matched competitor run or broad independent consensus. Start with one difficult task from your real workflow, compare the output and review time against the model you already use, and then decide whether the upgrade is worth the route-specific cost.



