Analysis · 2026-09-23 · 9 min read

Grok 4.7 API Pricing, 500K Context Window and Benchmarks

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens below 200K prompt tokens, with a 500K context window. Here is the official pricing, long-context tier, cache rate and xAI-published benchmark performance.

Grok 4.7 at a Glance

  • Official API price below 200K prompt tokens: $2.00 per 1M input, $0.50 per 1M cached input, and $6.00 per 1M output tokens.
  • Long-context price at 200K prompt tokens or more: $4.00 input, $1.00 cached input, and $12.00 output per 1M tokens. The higher rate applies to all tokens in that request.
  • Context window: 500,000 tokens.
  • Modalities: text and image input, with text output and no stated text-output limit.
  • Reasoning controls: low, medium, high (default), and xhigh.
  • Best suited to: coding, agentic tasks, long-running work and knowledge work.

!xAI

!OpenAI

!Anthropic

xAI released Grok 4.7 on September 21, 2026 as its new frontier model for coding and knowledge work. The official announcement says it uses a larger base model than Grok 4.6, received a longer reinforcement-learning run on harder multi-hour tasks, and was trained to verify its own work more carefully. This guide separates xAI's published facts from our cost analysis so buyers can compare the model without confusing vendor benchmarks with independent testing.

Official Grok 4.7 API Pricing

Grok 4.7 uses two price bands. The threshold is based on the number of prompt tokens, not the total context after generation.

Source: xAI API pricing and Grok 4.7 documentation, checked September 23, 2026.

The important detail is the 200K prompt cliff. Once a prompt reaches 200,000 tokens, xAI bills the request at the higher tier. It is not a marginal price that applies only to tokens above the threshold. A 199K-token prompt and a 200K-token prompt can therefore have meaningfully different bills.

What a Grok 4.7 request costs

At the standard tier, cached input is 75% cheaper than uncached input. For agents that repeatedly send the same system prompt, tool definitions or repository context, stable-prefix caching can matter as much as model choice.

Context Window and Technical Specs

A 500K context window is large enough for substantial codebases, contract collections and long research packets. Capacity is not the same as economy, however: the long-context price tier begins at 200K prompt tokens, well before the 500K ceiling. Retrieval, chunking and cache-aware prompt design remain valuable even when the model can technically accept the entire corpus.

Grok 4.7 accepts images alongside text, but returns text only. That makes it useful for screenshot analysis, document inspection and visual debugging without positioning it as an image-generation model.

Grok 4.7 Performance: xAI's Published Results

The following figures come from xAI's Grok 4.7 announcement and model card. They are vendor-reported results, not a TokenCost composite score or an independent retest. Test settings matter: xAI labels the DeepSWE result as high effort and presents Grok 4.7 at xhigh reasoning for the comparison.

Against Grok 4.6 in xAI's own table, Grok 4.7 improves from 40.4% to 46.3% on CursorBench 4.0, 53.0% to 64.0% on EEBench, and 20.3% to 38.0% on Terminal-Bench 4.0, while keeping the same headline $2/$6 standard-tier token price. That is the clearest performance-per-dollar story in the release.

These numbers should not be collapsed into a single universal rank. Coding, terminal use, legal work and clinical reasoning measure different capabilities, and reasoning effort can change both accuracy and token consumption. Use benchmark results to shortlist a model, then evaluate it on your own tasks and total completed-work cost.

Grok 4.7 vs Grok 4.6, GPT-5.6 Sol and Fable 5.1

This comparison reproduces the price and selected benchmark figures published in xAI's launch material. Competitor names and scores are xAI's labels and test results.

*xAI marks Grok 4.7's DeepSWE result as high effort.

Grok 4.7's launch positioning is not “best score everywhere.” It is frontier-level coding and agent performance at a lower token price than the premium models in xAI's comparison. The strongest buying case is therefore price-performance, especially for output-heavy workflows where its $6 per million output rate is materially below the $20 and $50 comparison rates.

Reasoning Effort and Real-World Cost

The API exposes four reasoning levels: low, medium, high and xhigh, with high as the default. Higher effort can improve difficult-task performance, but it may also produce more reasoning tokens and increase latency or cost. The economical deployment pattern is to route routine work to low or medium, then escalate only uncertain or high-value tasks.

For an agent, cost per successful task matters more than price per token:

1. Record input, cached input and output tokens for every attempt.

2. Track retries and tool-call failures by reasoning level.

3. Compare the total cost of a completed task, not just the first response.

4. Keep reusable instructions and tool schemas stable to maximize cache hits.

5. Trim or retrieve context before a prompt crosses the 200K price boundary.

Is Grok 4.7 Worth It?

Grok 4.7 is compelling for:

  • Coding agents that need long-horizon planning and verification.
  • Generation-heavy applications that benefit from a relatively low $6/M output rate.
  • Multimodal knowledge work using text and image inputs.
  • Large-context workflows that can keep repeated prefixes cached.

Watch the economics when:

  • Prompts regularly cross 200K tokens, doubling all three token rates.
  • xhigh reasoning produces long hidden reasoning traces for routine tasks.
  • A cheaper specialist model can complete classification, extraction or routing work reliably.

Bottom Line

Grok 4.7 combines a 500K context window, text-and-image input, adjustable reasoning and xAI-reported gains on coding, terminal, engineering and professional benchmarks. Its official standard-tier pricing is $2/M input, $0.50/M cached input and $6/M output. At 200K prompt tokens, those rates double to $4/$1/$12.

That price cliff is the operational detail to remember. Below it, Grok 4.7 is priced aggressively for a frontier coding and knowledge-work model. Above it, retrieval quality, context trimming and caching become essential. Check the live xAI pricing page, compare all models in the pricing table, and use the LLM leaderboard to follow new comparable benchmark coverage as it appears.

Official Sources

Related Reading