{"id":19000,"date":"2026-09-08T07:52:59","date_gmt":"2026-09-08T11:52:59","guid":{"rendered":"https:\/\/wp.glbgpt.com\/?p=19000"},"modified":"2026-09-08T08:10:19","modified_gmt":"2026-09-08T12:10:19","slug":"grok-4-api-pricing","status":"publish","type":"post","link":"https:\/\/wp.glbgpt.com\/it\/hub\/grok-4-api-pricing","title":{"rendered":"Grok 4 API: prezzi, costi, limiti e come iniziare"},"content":{"rendered":"<aside aria-label=\"Risposta rapida\" style=\"padding:22px;border-left:5px solid #17806b;background:#edf6f1;color:#1b362c;font:18px\/1.65 system-ui,sans-serif;margin:0 0 22px\"><strong style=\"display:block;font-size:15px;margin-bottom:8px;color:#126954\">Risposta rapida<\/strong>Grok 4 API pricing depends on the exact model ID. As of September 8, 2026, the retired <code>grok-4-0709<\/code> ID redirects to <code>grok-4.3<\/code>: <strong>$1.25 input, $0.20 cached input and $2.50 output per million tokens<\/strong> for prompts below 200K tokens. The current flagship, <code>grok-4.6<\/code>, costs <strong>$2.00 \/ $0.50 \/ $6.00<\/strong>, respectively. At 200K prompt tokens or more, those rates double. Reasoning is billed at the output rate. <a href=\"https:\/\/docs.x.ai\/developers\/pricing\">Official xAI pricing<\/a>.<\/aside>\n\n\n\n<p class=\"wp-block-paragraph\">Building an app or budgeting an agent? Start with the model you will actually run. The <a href=\"https:\/\/docs.x.ai\/developers\/migration\/may-15-retirement\">May 15, 2026 migration notice<\/a> says the original Grok 4 ID now resolves to Grok 4.3 with low reasoning effort. A successful call to an old ID no longer proves that the original model served it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">API usage is metered separately from Grok&#8217;s consumer subscriptions. <a href=\"https:\/\/www.glbgpt.com\/hub\/how-much-is-grok-4\/\">Grok subscription and access pricing<\/a> covers the chat plans; those monthly fees are not token credits for a direct xAI integration.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Want to try a prompt before wiring it into an app? <a href=\"https:\/\/www.glbgpt.com\/home\/grok-4-6?inviter=hub_grk46&amp;login=1\">Try Grok 4.6 on GlobalGPT<\/a> for a chat-based workflow. Use an xAI API key for your own integration and direct xAI billing. This article is published by GlobalGPT; product links use referral tracking.<\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<figure class=\"wp-block-image size-large\"><a href=\"https:\/\/www.glbgpt.com\/home\/grok-4-6?inviter=hub_grk46&amp;login=1\"><img alt=\"\" fetchpriority=\"high\" decoding=\"async\" width=\"1024\" height=\"581\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-1024x581.png\" class=\"wp-image-18146\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-1024x581.png 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-300x170.png 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-768x436.png 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-1536x872.png 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-2048x1162.png 2048w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/08\/image-15-18x10.png 18w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/figure>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-3e41869c wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/www.glbgpt.com\/home\/grok-4-6?inviter=hub_grk46&amp;login=1\">Prova subito Grok 4.6<\/a><\/div>\n<\/div>\n<\/div>\n\n\n\n<nav aria-label=\"Indice\" style=\"margin:28px 0;padding:24px;border-left:5px solid #a84b62;background:#f8edf0;color:#41232e;font:16px\/1.6 system-ui,sans-serif\"><strong style=\"display:block;font-size:20px;margin-bottom:12px\">Indice<\/strong><ol style=\"margin:0;padding-left:23px;display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,270px),1fr));gap:8px 28px\"><li><a href=\"#current-api-pricing\">Grok 4 API pricing: current rates and the retired original<\/a><\/li><li><a href=\"#token-billing\">How input, cached input and reasoning affect the bill<\/a><\/li><li><a href=\"#cost-examples\">Three transparent API cost examples<\/a><\/li><li><a href=\"#api-limits\">Context, output, rate limits and spending limits are different<\/a><\/li><li><a href=\"#real-usage\">Two real Grok 4.6 tasks: what the usage records reveal<\/a><\/li><li><a href=\"#developer-feedback\">What independent API users noticed<\/a><\/li><li><a href=\"#get-started\">How to get started with the official xAI API<\/a><\/li><li><a href=\"#errors-and-controls\">Common integration errors and cost controls<\/a><\/li><li><a href=\"#faq\">Grok 4 API pricing FAQ<\/a><\/li><\/ol><\/nav>\n\n\n\n<h2 id=\"current-api-pricing\" class=\"wp-block-heading\">Grok 4 API pricing: current rates and the retired original<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The table below uses <a href=\"https:\/\/docs.x.ai\/developers\/pricing\">xAI list prices checked on September 8, 2026<\/a>. Each amount is USD per one million tokens. Input, cached input and output are separate billable categories, so a single blended &#8220;price per million&#8221; hides the workload&#8217;s cost.<\/p>\n\n\n\n<section aria-label=\"Current API price table\" style=\"margin:24px 0;font:16px\/1.55 system-ui,sans-serif;color:#22352d\"><p style=\"margin:0 0 16px;font-weight:700\">USD per 1M tokens \u00b7 standard processing<\/p><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,260px),1fr));gap:16px\"><article style=\"min-width:0;padding:20px;border:1px solid #cddbd4;border-top:4px solid #197862;border-radius:6px\"><h3 style=\"font-size:22px;margin:0 0 4px\">grok-4.3<\/h3><p style=\"margin:0 0 16px;color:#4e6056\">Original Grok 4 redirect target<\/p><p style=\"margin:0 0 12px\"><strong>Contesto:<\/strong> 1.000.000 gettoni<\/p><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &lt; 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$1.25<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$0.20<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$2.50<\/dd><\/dl><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &gt;= 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$2.50<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$0.40<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$5.00<\/dd><\/dl><\/article><article style=\"min-width:0;padding:20px;border:1px solid #cddbd4;border-top:4px solid #ae5266;border-radius:6px\"><h3 style=\"font-size:22px;margin:0 0 4px\">grok-4.6<\/h3><p style=\"margin:0 0 16px;color:#4e6056\">Bandiera attuale<\/p><p style=\"margin:0 0 12px\"><strong>Contesto:<\/strong> 500,000 tokens<\/p><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &lt; 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$2.00<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$0.50<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$6.00<\/dd><\/dl><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &gt;= 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$4.00<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$1.00<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$12.00<\/dd><\/dl><\/article><article style=\"min-width:0;padding:20px;border:1px solid #cddbd4;border-top:4px solid #197862;border-radius:6px\"><h3 style=\"font-size:22px;margin:0 0 4px\">grok-4.5<\/h3><p style=\"margin:0 0 16px;color:#4e6056\">Earlier current-generation option<\/p><p style=\"margin:0 0 12px\"><strong>Contesto:<\/strong> 500,000 tokens<\/p><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &lt; 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$2.00<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$0.30<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$6.00<\/dd><\/dl><h4 style=\"font-size:15px;margin:18px 0 8px\">Prompt &gt;= 200K<\/h4><dl style=\"display:grid;grid-template-columns:1fr auto;gap:6px 12px;margin:0\"><dt>Ingresso<\/dt><dd style=\"margin:0;font-weight:700\">$4.00<\/dd><dt>Input memorizzato nella cache<\/dt><dd style=\"margin:0;font-weight:700\">$0.60<\/dd><dt>Output \/ reasoning<\/dt><dd style=\"margin:0;font-weight:700\">$12.00<\/dd><\/dl><\/article><\/div><p style=\"font-size:14px;color:#4d6156;margin:14px 0 0\">Once the prompt reaches 200K tokens, long-context rates apply to all tokens in that request, including output.<\/p><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The other listed Grok 4-family IDs are <code>grok-4.20-0309-ragionamento<\/code>, <code>grok-4.20-0309-senza ragionamento<\/code> e <code>grok-4.20-multi-agent-0309<\/code>. Each has a 1M-token context window and the same listed short\/long token rates as Grok 4.3. The matching unit price does not imply matching total token consumption, especially for multi-agent work.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a coding-specific alternative, <code>grok-build-0.1<\/code> has a 256K context window and short-context input\/cache\/output rates of $1.00\/$0.20\/$2.00, rising to $2.00\/$0.40\/$4.00 at the 200K prompt threshold. For broader behavior and use cases, see our <a href=\"https:\/\/www.glbgpt.com\/hub\/grok-4-6-review\/\">Recensione di Grok 4.6<\/a>.<\/p>\n\n\n\n<figure id=\"grok-api-media-catalog\" style=\"margin:28px 0\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/grok-official-model-catalog_2d8f5ebf7bf94fe3a8f48ad8c718010d.webp\"><img loading=\"lazy\" src=\"https:\/\/static.futureshareai.com\/glb_features\/grok-official-model-catalog_2d8f5ebf7bf94fe3a8f48ad8c718010d.webp\" width=\"1280\" height=\"581\" alt=\"Documentazione ufficiale di xAI che riporta il codice Grok 4.6, l&#039;ID del modello, un contesto di 500K e il ragionamento configurabile.\" decoding=\"async\" style=\"display:block;width:100%;height:auto;max-width:100%\"><\/a><figcaption style=\"font:14px\/1.6 system-ui,sans-serif;color:#4c5855;margin-top:10px\"><a href=\"https:\/\/docs.x.ai\/developers\/models\">Official xAI model catalog<\/a>, captured September 7 and rechecked September 8, 2026. The $2\/$6 headline is Grok 4.6 short-context input\/output pricing; cache and long-context rates are listed above.<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">What happened to the original $3 \/ $15 price?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Older references, including our <a href=\"https:\/\/www.glbgpt.com\/hub\/grok-4-api-guide\/\">Grok 4 API integration guide<\/a>, quote $3 per million input tokens and $15 per million output tokens for the original Grok 4. Treat that as a historical quote, not today&#8217;s bill for a retired ID. The old official model URL now leads to the general model catalog, so this is not a newly verified archival price sheet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Il <a href=\"https:\/\/x.ai\/news\/grok-4\">original Grok 4 announcement<\/a> describes a 256,000-token context window. That historical capacity, and older Fast-model prices or thresholds, must not be copied into a budget for the replacement. The precise rule for <code>grok-4-0709<\/code> is now Grok 4.3 with low reasoning effort, billed at Grok 4.3 rates.<\/p>\n\n\n\n<figure id=\"grok-api-media-retirement\" style=\"margin:28px 0\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/grok-api-retirement_931d106956cb4e78be50acac7166f538.webp\"><img loading=\"lazy\" src=\"https:\/\/static.futureshareai.com\/glb_features\/grok-api-retirement_931d106956cb4e78be50acac7166f538.webp\" width=\"1280\" height=\"784\" alt=\"Avviso ufficiale di ritiro che elenca grok-4-0709 e il relativo reindirizzamento a grok-4.3\" decoding=\"async\" style=\"display:block;width:100%;height:auto;max-width:100%\"><\/a><figcaption style=\"font:14px\/1.6 system-ui,sans-serif;color:#4c5855;margin-top:10px\"><a href=\"https:\/\/docs.x.ai\/developers\/migration\/may-15-retirement\">xAI retirement notice<\/a>, captured September 7, 2026; text rechecked September 8. The listed original Grok 4 ID redirects to Grok 4.3. The full notice separately identifies exceptions for coding and image models.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Grok 4 is a model generation, while Grok is also the name of a consumer product. Our <a href=\"https:\/\/www.glbgpt.com\/hub\/what-is-grok-4-the-ultimate-2026-full-guide-to-flagship-reasoning-model\/\">explanation of what Grok 4 is<\/a> provides that product background. For new code, explicitly select a supported model instead of relying on a retirement redirect or assuming a provider&#8217;s alias stays fixed.<\/p>\n\n\n\n<h2 id=\"token-billing\" class=\"wp-block-heading\">How input, cached input and reasoning affect the bill<\/h2>\n\n\n\n<section aria-label=\"One request, three token charges\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #167862;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">One request, three token charges<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Dati di input non memorizzati nella cache<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">(I &#8211; C) \u00d7 input rate<\/p><p style=\"margin:0\">I = all prompt tokens; C = cached prompt tokens.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Input memorizzato nella cache<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">C \u00d7 cache rate<\/p><p style=\"margin:0\">Cached tokens are part of input, not extra input.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Billable output<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">O \u00d7 output rate<\/p><p style=\"margin:0\">O includes reasoning when the endpoint reports it inside output.<\/p><\/div><\/div><\/section>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Token estimate in USD = [(I &#8211; C) \u00d7 input rate + C \u00d7 cached-input rate + O \u00d7 output rate] \/ 1,000,000.<\/strong> Then add applicable tool, storage or other charges. Select the long-context rates using the total prompt, including cached tokens, before doing the arithmetic.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Input includes the messages, instructions and conversation history sent to the model. Caching reuses an identical starting prefix across requests; it is automatic, and xAI recommends a conversation identifier to improve reuse. A similar-looking prompt does not guarantee a cache hit. The <a href=\"https:\/\/docs.x.ai\/developers\/advanced-api-usage\/prompt-caching\">prompt-caching documentation<\/a> explains the matching behavior.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In Chat Completions, read <code>prompt_tokens_details.cached_tokens<\/code>; in Responses, read <code>input_tokens_details.cached_tokens<\/code>. Il <a href=\"https:\/\/docs.x.ai\/developers\/advanced-api-usage\/prompt-caching\/usage-and-pricing\">cache usage and pricing reference<\/a> bills reasoning tokens at the full completion rate. If the output total already includes reasoning, adding <code>reasoning_tokens<\/code> again would double-count it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Grok 4.6 supports low, medium, high and xhigh reasoning effort; its documented default is high, and reasoning cannot be disabled. Grok 4.3 supports none, low, medium and high. That makes Grok 4.3 with none worth evaluating for strict extraction jobs, while Grok 4.6 may spend more tokens even when its visible answer is short. See the <a href=\"https:\/\/docs.x.ai\/developers\/model-capabilities\/text\/reasoning\">reasoning settings<\/a> and Grok 4.3&#8217;s model entry above.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Tool calls, batch work and priority requests<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Per the <a href=\"https:\/\/docs.x.ai\/developers\/pricing\">official tool price table<\/a>, Web Search, X Search and Code Execution each cost $5 per 1,000 calls, or $0.005 per call, plus token consumption. A request can invoke tools more than once. Ten search calls add $0.05 before model tokens; enabling a tool does not tell you how many times it will be used.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">File-attachment search costs $10 per 1,000 calls; Collections Search costs $2.50 per 1,000. Stored files and collections also have daily storage and download charges. Images and video use their own price units, so keep those out of a text-only token estimate. Our <a href=\"https:\/\/www.glbgpt.com\/hub\/how-to-generate-image-and-video-in-grok\/\">Grok image and video workflow tutorial<\/a> covers that separate generation route.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Il <a href=\"https:\/\/docs.x.ai\/developers\/advanced-api-usage\/batch-api\">API batch<\/a> offers a 20% token discount for Grok 4.3 and the three listed Grok 4.20 variants. The discount covers input, cached input, output and reasoning tokens. Grok 4.6&#8217;s model page marks Batch API as not supported. Do not apply a blanket 50% discount or assume every listed model supports batch processing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/docs.x.ai\/developers\/advanced-api-usage\/priority-processing\">Elaborazione prioritaria<\/a> costs 2x standard token rates, including cached and reasoning tokens, when the returned <code>service_tier<\/code> confirms <code>priority<\/code>. A fallback to the default tier uses standard rates. Priority applies to Chat Completions and Responses, not Batch API.<\/p>\n\n\n\n<h2 id=\"cost-examples\" class=\"wp-block-heading\">Three transparent API cost examples<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">These are list-price estimates, not observed invoices. All requests below use standard processing, text only, no tools, no storage, no retries and no taxes. &#8220;Output&#8221; includes any billable reasoning. The first two workloads assume prompts below 200K tokens.<\/p>\n\n\n\n<section aria-label=\"Cost per request and per 10,000 requests\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #a74d65;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">Cost per request and per 10,000 requests<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Grok 4.3 \u00b7 no cache<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">$0.005 \/ request<\/p><p style=\"margin:0\">2,000 input \u00d7 $1.25 + 1,000 output \u00d7 $2.50, divided by 1M. 10,000 requests = $50.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Grok 4.6 \u00b7 no cache<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">$0.010 \/ request<\/p><p style=\"margin:0\">2,000 input \u00d7 $2 + 1,000 output \u00d7 $6, divided by 1M. 10,000 requests = $100.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Grok 4.6 \u00b7 90% input cached<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">$0.0073 \/ request<\/p><p style=\"margin:0\">200 uncached \u00d7 $2 + 1,800 cached \u00d7 $0.50 + 1,000 output \u00d7 $6, divided by 1M. 10,000 requests = $73.<\/p><\/div><\/div><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">A high cache-hit rate helps most when input dominates the workload. In that Grok 4.6 example, caching saves $0.0027 per call, or 27% of the total, even though 90% of input tokens are cached. The output charge stays at $0.006. Judge savings against the entire request, not only the cached portion.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The 200K prompt threshold can change the whole request<\/h3>\n\n\n\n<section aria-label=\"One more input token crosses the price threshold\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #306b9c;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">One more input token crosses the price threshold<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">199,999 prompt tokens<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">$0.459998<\/p><p style=\"margin:0\">Grok 4.6: (199,999 \u00d7 $2 + 10,000 output \u00d7 $6) \/ 1M. No cache.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">200,000 prompt tokens<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">$0.920000<\/p><p style=\"margin:0\">Grok 4.6: (200,000 \u00d7 $4 + 10,000 output \u00d7 $12) \/ 1M. No cache.<\/p><\/div><\/div><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The main <a href=\"https:\/\/docs.x.ai\/developers\/pricing\">pricing table explicitly uses &#8220;Long context &gt;= 200k tokens&#8221;<\/a> and says all tokens use the higher rate once the prompt reaches the threshold. Some model-detail text says &#8220;exceed&#8221; 200K. For budgeting at exactly 200,000, use the explicit pricing-table boundary; do not assume the cheaper rate still applies.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a monthly forecast, estimate request volume separately for each workload, context band, cache-hit pattern and reasoning setting. Add expected retries and tool invocations. Then compare that estimate with your own returned cost data after a small pilot. A cents-per-call figure is only useful when its assumptions match production.<\/p>\n\n\n\n<h2 id=\"api-limits\" class=\"wp-block-heading\">Context, output, rate limits and spending limits are different<\/h2>\n\n\n\n<section aria-label=\"Four limits to check before scaling\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #306b9c;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">Four limits to check before scaling<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">Finestra contestuale<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">Per request<\/p><p style=\"margin:0\">Grok 4.3: 1M tokens. Grok 4.6: 500K. This is not a monthly allowance.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">Output budget<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">Per response<\/p><p style=\"margin:0\">Set a supported output cap. Context capacity does not guarantee that much visible output.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">RPS \/ TPM<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">Throughput<\/p><p style=\"margin:0\">Requests per second and tokens per minute. Both must stay within the team tier.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#306b9c;font-weight:700\">Spend controls<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">Money<\/p><p style=\"margin:0\">Prepaid balance, auto top-up and invoiced billing limits control separate funding behavior.<\/p><\/div><\/div><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">Il pubblico <a href=\"https:\/\/docs.x.ai\/developers\/rate-limits\">rate-limit table<\/a> lists Tier 0 at 37 requests per second and 10 million tokens per minute for Grok 4.3; Grok 4.6 is listed at 150 requests per second and 50 million tokens per minute. These are dated published tier entries, not a promise that every key has the same usable capacity. Your team&#8217;s Console is the authority for its current model access and limits.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Tiers are based on cumulative API spend since January 1, 2026: $0, $50, $250, $1,000 and $5,000 for Tiers 0 through 4. Requests are throttled per second using the per-minute request budget divided by 60, so you cannot spend a full minute&#8217;s allocation in one burst. Input, output, reasoning and cached tokens all count toward TPM, even when some cost less.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The checked Grok 4.3 and 4.6 model pages give context sizes without a separate hard maximum-output figure. The current <a href=\"https:\/\/docs.x.ai\/developers\/rest-api-reference\/inference\/responses\">Responses endpoint reference<\/a> defines <code>max_output_tokens<\/code> as the combined output-and-reasoning budget: it defaults to 128,000 when unset and accepts a larger value for longer generations. That default is not a guarantee of 128,000 visible answer tokens or an unlimited context allowance. Set an explicit budget and inspect incomplete responses.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/docs.x.ai\/console\/billing\">xAI billing documentation<\/a> separates prepaid credits from monthly invoiced billing, which is disabled by default. Auto top-up can replenish prepaid credits and has its own thresholds and monthly allowance. An invoiced billing limit controls postpaid usage; it is not the same thing as a single all-inclusive cap on prepaid purchases, top-ups and requests.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For your first integration, start with limited prepaid funding and review any auto top-up setting deliberately. For everyday chat access without writing an integration, the separate <a href=\"https:\/\/www.glbgpt.com\/hub\/how-to-use-grok-4\/\">how to use Grok guide<\/a> covers the product workflow. A chat subscription and an API team have different controls.<\/p>\n\n\n\n<h2 id=\"real-usage\" class=\"wp-block-heading\">Two real Grok 4.6 tasks: what the usage records reveal<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We reused two completed tests from September 7, 2026, through the Broly aggregation API at <code>https:\/\/anywhere.broly.ai\/v1<\/code>, using Chat Completions. Both requested and returned <code>grok-4.6<\/code>; the upstream model identity was not independently verified. These are Grok 4.6 gateway results, not original Grok 4 tests or a comparison across models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Each task ran once with <code>temperature=0<\/code>, <code>max_tokens=2048<\/code> e <code>stream=false<\/code>. No tools or explicit reasoning-effort setting were sent. The same system instruction was used for both requests:<\/p>\n\n\n\n<section id=\"system-prompt\" aria-label=\"Shared system instruction\" style=\"margin:22px 0;padding:20px;border-radius:6px;background:#1e2925;color:#f5faf6;font:15px\/1.6 system-ui,sans-serif;min-width:0\"><div style=\"display:flex;align-items:center;justify-content:space-between;gap:12px;flex-wrap:wrap;margin-bottom:12px\"><strong>Shared system instruction<\/strong><button type=\"button\" title=\"Copy Shared system instruction\" aria-label=\"Copy Shared system instruction\" style=\"min-width:44px;min-height:36px;padding:6px 12px;border:1px solid #7a9c8d;background:#d9ebe1;color:#173729;border-radius:4px;cursor:pointer\" onclick=\"(async(button)=>{const area=button.closest(&#8216;section&#8217;).querySelector(&#8216;textarea&#8217;);const status=button.closest(&#8216;section&#8217;).querySelector(&#8216;[role=status]&#8217;);try{if(!navigator.clipboard)throw Error();await Promise.race([navigator.clipboard.writeText(area.value),new Promise((_,reject)=>setTimeout(reject,700))]);status.textContent=&#8217;Copied&#8217;;}catch{area.focus();area.select();status.textContent=&#8217;Press Ctrl\/Cmd+C&#8217;;}})(this)&#8221;>Copy<\/button><\/div><textarea readonly aria-label=\"shared system instruction text\">Complete the user&#8217;s self-contained task using only the supplied information. Follow the requested output format exactly. Do not claim to have used external tools.<\/textarea><span role=\"status\" aria-live=\"polite\" style=\"display:block;min-height:24px;font-size:13px;margin-top:7px\"><\/span><\/section>\n\n\n\n<h3 class=\"wp-block-heading\">Expense extraction<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The task converted five synthetic expense notes into ordered JSON with integer cents, including a negative refund. The response preserved every requested field and produced the correct $209.00 net total.<\/p>\n\n\n\n<section id=\"prompt-expense-extraction\" aria-label=\"Exact prompt: Expense extraction\" style=\"margin:22px 0;padding:20px;border-radius:6px;background:#1e2925;color:#f5faf6;font:15px\/1.6 system-ui,sans-serif;min-width:0\"><div style=\"display:flex;align-items:center;justify-content:space-between;gap:12px;flex-wrap:wrap;margin-bottom:12px\"><strong>Exact prompt: Expense extraction<\/strong><button type=\"button\" title=\"Copy Exact prompt: Expense extraction\" aria-label=\"Copy Exact prompt: Expense extraction\" style=\"min-width:44px;min-height:36px;padding:6px 12px;border:1px solid #7a9c8d;background:#d9ebe1;color:#173729;border-radius:4px;cursor:pointer\" onclick=\"(async(button)=>{const area=button.closest(&#8216;section&#8217;).querySelector(&#8216;textarea&#8217;);const status=button.closest(&#8216;section&#8217;).querySelector(&#8216;[role=status]&#8217;);try{if(!navigator.clipboard)throw Error();await Promise.race([navigator.clipboard.writeText(area.value),new Promise((_,reject)=>setTimeout(reject,700))]);status.textContent=&#8217;Copied&#8217;;}catch{area.focus();area.select();status.textContent=&#8217;Press Ctrl\/Cmd+C&#8217;;}})(this)&#8221;>Copy<\/button><\/div><textarea readonly aria-label=\"exact prompt: expense extraction text\">Normalize these synthetic USD expense notes. Return only valid JSON, without Markdown or extra text, with exactly these top-level keys: currency, items, net_total_cents. currency must be USD. Each item must have exactly label and amount_cents. Preserve note order. Amounts must be integer cents; a refund is negative. Do not invent additional expenses.\n\nNotes:\nTrain ticket: $46.80\nHotel: $128.00\nDinner: $27.45\nTicket refund received: $12.50\nTaxi: $19.25\n\nUse these labels respectively: train, hotel, dinner, refund, taxi.<\/textarea><span role=\"status\" aria-live=\"polite\" style=\"display:block;min-height:24px;font-size:13px;margin-top:7px\"><\/span><\/section>\n\n\n\n<section aria-label=\"Complete returned output: Expense extraction\" style=\"margin:20px 0;padding:18px;border-left:4px solid #258773;background:#edf5f1;color:#173729\"><h4 style=\"font:700 16px\/1.5 system-ui,sans-serif;margin:0 0 12px\">Complete returned output: Expense extraction<\/h4><pre style=\"margin:0;white-space:pre-wrap;overflow-wrap:anywhere;font:14px\/1.65 Consolas,monospace\">{\n  &quot;currency&quot;: &quot;USD&quot;,\n  &quot;items&quot;: [\n    {\n      &quot;label&quot;: &quot;train&quot;,\n      &quot;amount_cents&quot;: 4680\n    },\n    {\n      &quot;label&quot;: &quot;hotel&quot;,\n      &quot;amount_cents&quot;: 12800\n    },\n    {\n      &quot;label&quot;: &quot;dinner&quot;,\n      &quot;amount_cents&quot;: 2745\n    },\n    {\n      &quot;label&quot;: &quot;refund&quot;,\n      &quot;amount_cents&quot;: -1250\n    },\n    {\n      &quot;label&quot;: &quot;taxi&quot;,\n      &quot;amount_cents&quot;: 1925\n    }\n  ],\n  &quot;net_total_cents&quot;: 20900\n}<\/pre><\/section>\n\n\n\n<section aria-label=\"Reported usage \u00b7 Expense extraction\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #167862;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">Reported usage \u00b7 Expense extraction<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Input \/ cached<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">364 \/ 192<\/p><p style=\"margin:0\">Cached tokens are included in the input count.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Completion \/ reasoning<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">991 \/ 857<\/p><p style=\"margin:0\">Reasoning is reported inside completion details.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#167862;font-weight:700\">Wall-clock time<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">22.753 seconds<\/p><p style=\"margin:0\">Includes network and gateway overhead; one run.<\/p><\/div><\/div><\/section>\n\n\n\n<h3 class=\"wp-block-heading\">Constrained scheduling<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The task selected indivisible jobs under an eight-hour budget with prerequisites. The response chose A, B, C and E, totaling eight hours and value 24. We independently enumerated all 64 subsets: 17 were feasible, and this was the unique optimal set.<\/p>\n\n\n\n<section id=\"prompt-constrained-scheduling\" aria-label=\"Exact prompt: Constrained scheduling\" style=\"margin:22px 0;padding:20px;border-radius:6px;background:#1e2925;color:#f5faf6;font:15px\/1.6 system-ui,sans-serif;min-width:0\"><div style=\"display:flex;align-items:center;justify-content:space-between;gap:12px;flex-wrap:wrap;margin-bottom:12px\"><strong>Exact prompt: Constrained scheduling<\/strong><button type=\"button\" title=\"Copy Exact prompt: Constrained scheduling\" aria-label=\"Copy Exact prompt: Constrained scheduling\" style=\"min-width:44px;min-height:36px;padding:6px 12px;border:1px solid #7a9c8d;background:#d9ebe1;color:#173729;border-radius:4px;cursor:pointer\" onclick=\"(async(button)=>{const area=button.closest(&#8216;section&#8217;).querySelector(&#8216;textarea&#8217;);const status=button.closest(&#8216;section&#8217;).querySelector(&#8216;[role=status]&#8217;);try{if(!navigator.clipboard)throw Error();await Promise.race([navigator.clipboard.writeText(area.value),new Promise((_,reject)=>setTimeout(reject,700))]);status.textContent=&#8217;Copied&#8217;;}catch{area.focus();area.select();status.textContent=&#8217;Press Ctrl\/Cmd+C&#8217;;}})(this)&#8221;>Copy<\/button><\/div><textarea readonly aria-label=\"exact prompt: constrained scheduling text\">Choose the highest-value feasible set of these synthetic jobs for one worker with an 8-hour budget. Jobs are indivisible, cannot overlap, and a prerequisite job must be selected and finished before its dependent job starts. Unused time is allowed. Each job can be performed at most once.\n\nJob A: 3 hours, value 8, no prerequisites.\nJob B: 2 hours, value 7, prerequisite A.\nJob C: 2 hours, value 6, no prerequisites.\nJob D: 4 hours, value 12, prerequisite C.\nJob E: 1 hour, value 3, no prerequisites.\nJob F: 3 hours, value 10, prerequisite E.\n\nReturn only valid JSON without Markdown, with exactly these keys: selected, order, total_hours, total_value. selected must list the selected job IDs alphabetically. order must be a valid chronological order for those jobs. Do not include a reasoning narrative.<\/textarea><span role=\"status\" aria-live=\"polite\" style=\"display:block;min-height:24px;font-size:13px;margin-top:7px\"><\/span><\/section>\n\n\n\n<section aria-label=\"Complete returned output: Constrained scheduling\" style=\"margin:20px 0;padding:18px;border-left:4px solid #258773;background:#edf5f1;color:#173729\"><h4 style=\"font:700 16px\/1.5 system-ui,sans-serif;margin:0 0 12px\">Complete returned output: Constrained scheduling<\/h4><pre style=\"margin:0;white-space:pre-wrap;overflow-wrap:anywhere;font:14px\/1.65 Consolas,monospace\">{\n  &quot;selected&quot;: [&quot;A&quot;, &quot;B&quot;, &quot;C&quot;, &quot;E&quot;],\n  &quot;order&quot;: [&quot;A&quot;, &quot;B&quot;, &quot;C&quot;, &quot;E&quot;],\n  &quot;total_hours&quot;: 8,\n  &quot;total_value&quot;: 24\n}<\/pre><\/section>\n\n\n\n<section aria-label=\"Reported usage \u00b7 Constrained scheduling\" style=\"margin:24px 0;padding:22px 0;border-top:3px solid #a74d65;border-bottom:1px solid #d4dfda;font:16px\/1.55 system-ui,sans-serif;color:#23332d\"><h3 style=\"font-size:21px;line-height:1.35;margin:0 0 18px\">Reported usage \u00b7 Constrained scheduling<\/h3><div style=\"display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,220px),1fr));gap:22px\"><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Input \/ cached<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">428 \/ 192<\/p><p style=\"margin:0\">Cached tokens are included in the input count.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Completion \/ reasoning<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">4267 \/ 4217<\/p><p style=\"margin:0\">Reasoning is reported inside completion details.<\/p><\/div><div style=\"min-width:0\"><p style=\"margin:0 0 5px;color:#a74d65;font-weight:700\">Wall-clock time<\/p><p style=\"font-size:26px;font-weight:700;line-height:1.25;margin:0 0 9px;overflow-wrap:anywhere\">61.980 seconds<\/p><p style=\"margin:0\">Includes network and gateway overhead; one run.<\/p><\/div><\/div><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The scheduling answer is shorter than the expense JSON, yet its reported completion total is 4,267 tokens, including 4,217 reasoning tokens. That is the budgeting lesson: visible answer length is a poor proxy for reasoning-model consumption. These two tasks are useful examples of usage shape, not a latency benchmark or a general quality ranking.<\/p>\n\n\n\n<section aria-label=\"Gateway usage excerpt: constrained scheduling\" style=\"margin:20px 0;padding:18px;border-left:4px solid #258773;background:#edf5f1;color:#173729\"><h4 style=\"font:700 16px\/1.5 system-ui,sans-serif;margin:0 0 12px\">Gateway usage excerpt: constrained scheduling<\/h4><pre style=\"margin:0;white-space:pre-wrap;overflow-wrap:anywhere;font:14px\/1.65 Consolas,monospace\">{\n  &quot;prompt_tokens&quot;: 428,\n  &quot;completion_tokens&quot;: 4267,\n  &quot;total_tokens&quot;: 4695,\n  &quot;prompt_tokens_details&quot;: {\n    &quot;cached_tokens&quot;: 192\n  },\n  &quot;completion_tokens_details&quot;: {\n    &quot;reasoning_tokens&quot;: 4217\n  }\n}<\/pre><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The gateway also returned zero-valued <code>input_tokens<\/code> e <code>output_tokens<\/code> fields, despite nonzero Chat Completions counters. For these records, we used <code>prompt_tokens<\/code> e <code>completion_tokens<\/code>. We did not treat the zero aliases as a free request or infer visible-token counts from other zero fields.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Applying official rates without calling it an invoice<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">If these gateway counts were billed using Grok 4.6&#8217;s current short-context xAI list prices, the extraction task would be <strong>[(364 &#8211; 192) \u00d7 $2 + 192 \u00d7 $0.50 + 991 \u00d7 $6] \/ 1M = $0.006386<\/strong>. Scheduling would be <strong>[(428 &#8211; 192) \u00d7 $2 + 192 \u00d7 $0.50 + 4,267 \u00d7 $6] \/ 1M = $0.026170<\/strong>. These are hypothetical xAI-rate estimates, not charges observed at Broly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The scheduling response reported 4,267 completion tokens even though the request sent <code>max_tokens=2048<\/code>. The gateway&#8217;s cap semantics were not verified. This discrepancy does not establish how xAI&#8217;s own output cap works, and the requested cap cannot be treated as a proven spending ceiling for this gateway.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Direct xAI responses have a more useful billing signal: <a href=\"https:\/\/docs.x.ai\/developers\/cost-tracking\">cost_in_usd_ticks<\/a>. The official documentation defines one USD as 10,000,000,000 ticks and says the field includes that request&#8217;s token charges, discounts and server-side tool calls. The reused gateway responses do not contain that field. For a conversation, sum the per-request tick values rather than assuming each response contains a session total.<\/p>\n\n\n\n<h2 id=\"developer-feedback\" class=\"wp-block-heading\">What independent API users noticed<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">On August 21, 2026, Hacker News user <a href=\"https:\/\/news.ycombinator.com\/item?id=49394980\">mediaman<\/a> pointed out that Grok 4.6&#8217;s $0.50-per-million cached-input rate is 25% of its $2 uncached-input rate. Their concern was repeated cached reads in agentic work. The two Grok rates match the current official table; the commenter&#8217;s wider comparisons and speculation are personal assessments.<\/p>\n\n\n\n<figure id=\"grok-api-media-developer-review-1\" style=\"margin:28px 0\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/developer-review-1_345c99350c9941658de7b0d2b6ebb8a3.webp\"><img loading=\"lazy\" src=\"https:\/\/static.futureshareai.com\/glb_features\/developer-review-1_345c99350c9941658de7b0d2b6ebb8a3.webp\" width=\"1280\" height=\"226\" alt=\"mediaman on Hacker News discussing Grok API cache pricing\" decoding=\"async\" style=\"display:block;width:100%;height:auto;max-width:100%\"><\/a><figcaption style=\"font:14px\/1.6 system-ui,sans-serif;color:#4c5855;margin-top:10px\">mediaman on Grok 4.6 cache pricing, August 21, 2026. Personal cost assessment; comparisons and speculation are the commenter&#8217;s. Captured September 8, 2026.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Earlier, on November 9, 2025, <a href=\"https:\/\/news.ycombinator.com\/item?id=45870486\">XCSme<\/a> described using Grok 4 Fast via API for data parsing and extraction, calling it cheap and fast. That is a historical report about a now-retired Fast model. It does not establish today&#8217;s price, speed or quality for Grok 4.3 or Grok 4.6.<\/p>\n\n\n\n<figure id=\"grok-api-media-developer-review-2\" style=\"margin:28px 0\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/developer-review-2_00910b83138c4d11bd189750241e31ff.webp\"><img loading=\"lazy\" src=\"https:\/\/static.futureshareai.com\/glb_features\/developer-review-2_00910b83138c4d11bd189750241e31ff.webp\" width=\"1280\" height=\"125\" alt=\"XCSme on Hacker News discussing Grok API data extraction\" decoding=\"async\" style=\"display:block;width:100%;height:auto;max-width:100%\"><\/a><figcaption style=\"font:14px\/1.6 system-ui,sans-serif;color:#4c5855;margin-top:10px\">XCSme on using Grok 4 Fast for extraction, November 9, 2025. Historical personal experience with a now-retired model. Captured September 8, 2026.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">These two individual experiences suggest useful pilot questions: how much of your bill comes from cached context, and how much from reasoning and retries? They do not establish a developer consensus. If you are also comparing the resulting work, our <a href=\"https:\/\/www.glbgpt.com\/hub\/grok-vs-chatgpt-which-one-actually-fits-your-work\/\">Grok versus ChatGPT work, search and coding comparison<\/a> addresses that separate decision.<\/p>\n\n\n\n<h2 id=\"get-started\" class=\"wp-block-heading\">How to get started with the official xAI API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Il <a href=\"https:\/\/docs.x.ai\/developers\/quickstart\">official xAI quickstart<\/a> uses the Responses API at <code>https:\/\/api.x.ai\/v1\/responses<\/code>. This is a different endpoint and account from the aggregation venue used in our September 7 tests.<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Create or sign in to your xAI Console account and select the correct team.<\/li>\n\n\n\n<li>Add API credits and review spending settings. Confirm the selected model is available to your team.<\/li>\n\n\n\n<li>Create an API key and supply it to your server process as XAI_API_KEY. Keep it out of browser code, source control, screenshots and logs.<\/li>\n\n\n\n<li>Send a small request using an explicit current model ID, then record the returned model, usage, service tier and request identifier.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The example below uses Node.js 18 or later and its built-in fetch, so no SDK installation is required. Set XAI_API_KEY in your process environment before running it. It uses the current official Responses request shape and low reasoning effort; the code was checked locally without sending another paid API request.<\/p>\n\n\n\n<section id=\"first-request\" aria-label=\"First official request \u00b7 Node.js\" style=\"margin:22px 0;padding:20px;border-radius:6px;background:#1e2925;color:#f5faf6;font:15px\/1.6 system-ui,sans-serif;min-width:0\"><div style=\"display:flex;align-items:center;justify-content:space-between;gap:12px;flex-wrap:wrap;margin-bottom:12px\"><strong>First official request \u00b7 Node.js<\/strong><button type=\"button\" title=\"Copy First official request \u00b7 Node.js\" aria-label=\"Copy First official request \u00b7 Node.js\" style=\"min-width:44px;min-height:36px;padding:6px 12px;border:1px solid #7a9c8d;background:#d9ebe1;color:#173729;border-radius:4px;cursor:pointer\" onclick=\"(async(button)=>{const area=button.closest(&#8216;section&#8217;).querySelector(&#8216;textarea&#8217;);const status=button.closest(&#8216;section&#8217;).querySelector(&#8216;[role=status]&#8217;);try{if(!navigator.clipboard)throw Error();await Promise.race([navigator.clipboard.writeText(area.value),new Promise((_,reject)=>setTimeout(reject,700))]);status.textContent=&#8217;Copied&#8217;;}catch{area.focus();area.select();status.textContent=&#8217;Press Ctrl\/Cmd+C&#8217;;}})(this)&#8221;>Copy<\/button><\/div><textarea readonly aria-label=\"first official request \u00b7 node.js text\">const apiKey = process.env.XAI_API_KEY;\nif (!apiKey) throw new Error(&#8216;Set XAI_API_KEY first&#8217;);\n\nconst response = await fetch(&#8216;https:\/\/api.x.ai\/v1\/responses&#8217;, {\n  method: &#8216;POST&#8217;,\n  headers: {\n    Authorization: &#8216;Bearer &#8216; + apiKey,\n    &#8216;Content-Type&#8217;: &#8216;application\/json&#8217;,\n  },\n  body: JSON.stringify({\n    model: &#8216;grok-4.6&#8217;,\n    reasoning: { effort: &#8216;low&#8217; },\n    max_output_tokens: 2048,\n    input: &#8216;Return only the sum of 12 and 30.&#8217;,\n  }),\n  signal: AbortSignal.timeout(120_000),\n});\n\nif (!response.ok) {\n  throw new Error(&#8216;xAI HTTP &#8216; + response.status);\n}\nconst data = await response.json();\nconst text = (data.output ?? [])\n  .filter(item =&gt; item.type === &#8216;message&#8217;)\n  .flatMap(item =&gt; item.content ?? [])\n  .filter(item =&gt; item.type === &#8216;output_text&#8217;)\n  .map(item =&gt; item.text)\n  .join(&#8216;\\n&#8217;);\n\nconsole.log(text);\nconsole.log(&#8216;Status:&#8217;, data.status);\nconsole.log(&#8216;Model:&#8217;, data.model);\nconsole.log(&#8216;Service tier:&#8217;, data.service_tier);\nconsole.log(&#8216;Usage:&#8217;, data.usage);\nconst ticks = data.usage?.cost_in_usd_ticks;\nif (typeof ticks === &#8216;number&#8217;) {\n  console.log(&#8216;Request cost USD:&#8217;, ticks \/ 1e10);\n} else {\n  console.log(&#8216;Request cost field not returned&#8217;);\n}<\/textarea><span role=\"status\" aria-live=\"polite\" style=\"display:block;min-height:24px;font-size:13px;margin-top:7px\"><\/span><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">Look for a completed status and inspect the usage object before scaling up. This example budgets 2,048 output-and-reasoning tokens through the official Responses field. A timeout is a client-side deadline, not proof that server-side work was cancelled or unbilled. The example does not retry automatically; the token cap is not an all-inclusive dollar cap on an application that also sends input or invokes tools.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For an OpenAI-compatible SDK, the base URL is <code>https:\/\/api.x.ai\/v1<\/code> and the credential is still your <strong>xAI<\/strong> API key. The Responses API uses <code>input<\/code> e <code>reasoning: { effort: \"low\" }<\/code>; Chat Completions uses <code>messaggi<\/code> and its own parameter names. Copying a gateway request unchanged into a different endpoint can fail even when both accept JSON.<\/p>\n\n\n\n<h2 id=\"errors-and-controls\" class=\"wp-block-heading\">Common integration errors and cost controls<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/docs.x.ai\/developers\/debugging\">401 Non autorizzato<\/a>: check the Bearer header, key and process environment. A 403 can indicate team permissions or access restrictions. Repeating the same invalid request does not repair it.<\/li>\n\n\n\n<li>404 Not Found: inspect the exact endpoint and model ID. Distinguish an invalid ID from a documented retired ID that still resolves through a redirect.<\/li>\n\n\n\n<li>400 or 422: validate the request schema and supported settings. For example, the current reasoning documentation rejects presence\/frequency penalties and stop settings for reasoning models.<\/li>\n\n\n\n<li>429 Too Many Requests: queue work, smooth bursts and use bounded exponential backoff with jitter. Track both RPS and TPM, and include retries in the workload budget.<\/li>\n\n\n\n<li>Unexpected output or incomplete status: inspect the response, reasoning setting and applicable output cap. Validate machine-readable output before consuming it; a successful HTTP response is not a task-quality guarantee.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">For cost control, choose the least expensive model and reasoning setting that passes your own task checks, keep stable instructions at the start for cache reuse, remove irrelevant history, monitor the 200K prompt boundary and limit agent tool loops. Use batch only when that model supports it and the work can wait. Treat priority as a separate purchasing choice.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Keep token estimates for planning and returned cost ticks for reconciliation. Store integer ticks when aggregating many requests, and compare request totals with Console usage and invoices. A missing cost field means &#8220;unknown,&#8221; not zero; a small visible reply does not prove a small bill.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For Grok 4 API pricing today, build your budget around the model that actually serves the request, the prompt&#8217;s price band and total billable output. Start with a small official request, inspect its usage and cost field, then scale the workload that passes your own checks. For prompt exploration in a chat workspace, <a href=\"https:\/\/www.glbgpt.com\/home\/grok-4-6?inviter=hub_grk46&amp;login=1\">open Grok 4.6 on GlobalGPT<\/a>.<\/p>\n\n\n\n<h2 id=\"faq\" class=\"wp-block-heading\">Grok 4 API pricing FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">How much does the Grok 4 API cost now?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">As of September 8, 2026, the retired grok-4-0709 ID redirects to Grok 4.3. Below 200K prompt tokens, its USD-per-million rates are $1.25 input, $0.20 cached input and $2.50 output. Long-context rates double.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Is the original Grok 4 API still available?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The original grok-4-0709 model was retired from the xAI API on May 15, 2026. Its ID continues to resolve, but requests are served by Grok 4.3 with low reasoning effort and billed at the replacement rates.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How much is the Grok 4.6 API?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Grok 4.6 costs $2 input, $0.50 cached input and $6 output per million tokens below 200K prompt tokens. At 200K or more, the rates are $4, $1 and $12. Tool calls and other applicable charges are separate.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Are reasoning tokens charged?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. xAI bills reasoning tokens at the completion rate. When reasoning is already included in the reported output or completion total, do not add it a second time in your estimate.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does cached input count toward the 200K threshold?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. Count the full prompt, including cached tokens. The official pricing table puts a prompt of 200K tokens or more in the long-context band, and the higher rates apply to all tokens in that request.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Is there a free Grok API allowance with a chat subscription?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Do not assume a chat subscription includes xAI API credits. The official API quickstart uses a separate Console account and credits. Any promotion must be verified for that API team; consumer free access is a different offer.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What are the Grok API rate limits?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Limits depend on the model and team tier. The checked Tier 0 table lists Grok 4.3 at 37 RPS and 10M TPM, and Grok 4.6 at 150 RPS and 50M TPM. Use your team Console to confirm the limits that apply to your key.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Posso usare il modello Grok 4.6 sul modello GlobalGPT?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. GlobalGPT has a verified Grok 4.6 chat entry. That is a separate platform workflow; it does not replace an xAI API key, direct API billing or the controls in your xAI Console.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How do I check the actual cost of an xAI request?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Read usage.cost_in_usd_ticks from the official response and divide by 10,000,000,000 for USD. The field covers that request, including token discounts and server-side tools. Sum requests to track a conversation.<\/p>\n\n\n\n<script type=\"application\/ld+json\">{\n    \"@context\": \"https:\\\/\\\/schema.org\",\n    \"@type\": \"FAQPage\",\n    \"mainEntity\": [\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How much does the Grok 4 API cost now?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"As of September 8, 2026, the retired grok-4-0709 ID redirects to Grok 4.3. Below 200K prompt tokens, its USD-per-million rates are $1.25 input, $0.20 cached input and $2.50 output. Long-context rates double.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Is the original Grok 4 API still available?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"The original grok-4-0709 model was retired from the xAI API on May 15, 2026. Its ID continues to resolve, but requests are served by Grok 4.3 with low reasoning effort and billed at the replacement rates.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How much is the Grok 4.6 API?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Grok 4.6 costs $2 input, $0.50 cached input and $6 output per million tokens below 200K prompt tokens. At 200K or more, the rates are $4, $1 and $12. Tool calls and other applicable charges are separate.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Are reasoning tokens charged?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. xAI bills reasoning tokens at the completion rate. When reasoning is already included in the reported output or completion total, do not add it a second time in your estimate.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Does cached input count toward the 200K threshold?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. Count the full prompt, including cached tokens. The official pricing table puts a prompt of 200K tokens or more in the long-context band, and the higher rates apply to all tokens in that request.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Is there a free Grok API allowance with a chat subscription?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Do not assume a chat subscription includes xAI API credits. The official API quickstart uses a separate Console account and credits. Any promotion must be verified for that API team; consumer free access is a different offer.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"What are the Grok API rate limits?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Limits depend on the model and team tier. The checked Tier 0 table lists Grok 4.3 at 37 RPS and 10M TPM, and Grok 4.6 at 150 RPS and 50M TPM. Use your team Console to confirm the limits that apply to your key.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can I use Grok 4.6 on GlobalGPT?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. GlobalGPT has a verified Grok 4.6 chat entry. That is a separate platform workflow; it does not replace an xAI API key, direct API billing or the controls in your xAI Console.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"How do I check the actual cost of an xAI request?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Read usage.cost_in_usd_ticks from the official response and divide by 10,000,000,000 for USD. The field covers that request, including token discounts and server-side tools. Sum requests to track a conversation.\"\n            }\n        }\n    ]\n}<\/script>","protected":false},"excerpt":{"rendered":"<p>Quick AnswerGrok 4 API pricing depends on the exact model ID. As of September 8, 2026, the retired grok-4-0709 ID redirects to grok-4.3: $1.25 input, $0.20 cached input and $2.50 output per million tokens for prompts below 200K tokens. The current flagship, grok-4.6, costs $2.00 \/ $0.50 \/ $6.00, respectively. At 200K prompt tokens or [&hellip;]<\/p>","protected":false},"author":16,"featured_media":19009,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_seopress_robots_primary_cat":"","_seopress_titles_title":"Grok 4 API Pricing: Token Costs, Limits & Setup","_seopress_titles_desc":"Budget Grok 4 API pricing with real usage records and worked cost examples. Spot reasoning charges and the 200K price jump before you scale your app.","_seopress_robots_index":"","footnotes":""},"categories":[7],"tags":[],"class_list":["post-19000","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-chat"],"_links":{"self":[{"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/posts\/19000","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/comments?post=19000"}],"version-history":[{"count":2,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/posts\/19000\/revisions"}],"predecessor-version":[{"id":19005,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/posts\/19000\/revisions\/19005"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/media\/19009"}],"wp:attachment":[{"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/media?parent=19000"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/categories?post=19000"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/it\/wp-json\/wp\/v2\/tags?post=19000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}