Grok 4.7 vs 4.6: Price, Context and Upgrade Guide
Compare current API rates, context limits and evidence-based evaluation results for Grok 4.7 and 4.6, then choose an upgrade path.
This comparison was checked on September 26, 2026. Customer-facing prices come from the public BazaarLink model catalogue; capability details come from xAI documentation and announcement material.
Price and context
| Model and prompt tier | Input | Output | Context |
|---|---|---|---|
| Grok 4.7, up to 200K prompt tokens | $1.60 / 1M | $4.80 / 1M | 500,000 |
| Grok 4.7, above 200K prompt tokens | $3.20 / 1M | $9.60 / 1M | 500,000 |
| Grok 4.6, up to 200K prompt tokens | $2.00 / 1M | $6.00 / 1M | 500,000 |
| Grok 4.6, above 200K prompt tokens | $4.00 / 1M | $12.00 / 1M | 500,000 |
At both published prompt tiers, Grok 4.7 is 20% lower for input and output in the current BazaarLink catalogue. The context limit is the same 500,000 tokens for both rows in the catalogue.
What can be reproduced
xAI’s Grok 4.7 announcement describes a larger base model, longer reinforcement learning and improved self-verification. Treat those as the maker’s description of the release, not as a universal outcome for every prompt.
Independent public comparisons are useful but need careful reading. The Model Gap reports HLE 43.1 for Grok 4.7 xhigh and 42.9 for Grok 4.6 high, while LiveBench is 77.4 versus 78.0. The effort settings differ, so those numbers do not establish a same-setting winner. A BitsMinds review cites Artificial Analysis values of 46 for Grok 4.7 xhigh and 44 for Grok 4.6 high; that comparison also uses different effort settings.
To run a reproducible comparison, freeze a prompt set, keep the same answer schema, use the same effort setting where both models support it, record the model ID, and score correctness and format adherence separately. Store the input and output token counts with each result so quality and cost can be reviewed together. Do not infer a general winner from a small sample.
Which model should you choose?
| Situation | Starting choice | Why |
|---|---|---|
| You want to test the newer release and the current price delta is acceptable | Grok 4.7 | It is the newer xAI release and currently has the lower catalogue rate |
| Your existing prompts are stable and the older model meets the quality bar | Grok 4.6 | Keep the known behavior until a fixed-sample comparison shows a benefit |
| You need a defensible migration decision | Run both first | Compare the same prompts, schema, effort and token mix |
The practical upgrade question is not whether a model label sounds newer. It is whether Grok 4.7 improves your measured task set enough to justify changing prompts, validation and operating cost.
Sources
FAQ
Is Grok 4.7 cheaper than Grok 4.6?
Yes, in the current BazaarLink catalogue. Input and output are each 20% lower at both listed prompt tiers: up to 200K and above 200K prompt tokens.
Do Grok 4.7 and 4.6 have the same context limit?
The current catalogue lists a 500,000-token context limit for both models.
Do public benchmarks prove that Grok 4.7 is always better?
No. Public comparisons use different datasets and effort settings. Use a fixed, representative prompt set and score both models under the same conditions.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API