Grok 4.7 Held Its Price and Doubled Its Token Bill
xAI's new model keeps Grok 4.6's $2/$6 pricing and posts real gains, but independent testing shows it writes more than twice as many tokens per task, which changes what it actually costs.
xAI released Grok 4.7 on Monday, September 21, and called it "our most capable model for coding and knowledge work." The price did not move. Grok 4.7 costs $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens below 200,000 prompt tokens, the same as Grok 4.6, according to xAI's API release notes. Requests above that threshold cost double.
The benchmark gains are real. Independent testers confirm most of them. The catch is in a number xAI did not put on its launch page: how many tokens the model uses to finish a task. According to Artificial Analysis, Grok 4.7 at its highest reasoning setting produced about 81,000 output tokens per Intelligence Index task, more than double the roughly 36,000 Grok 4.6 used. A per-token price that stays flat does not mean a per-task cost that stays flat.
What xAI shipped
xAI says Grok 4.7 uses a new, larger base model, a longer reinforcement learning run, and training weighted toward difficult problems that take hours to solve. It also claims better self-verification and better handling of long context. Decrypt reports the model has 2.1 trillion parameters, up from 1.5 trillion for Grok 4.6, and was trained with supplemental SpaceX data including Starlink telemetry, manufacturing records and engineering failure logs. xAI's own announcement does not give a parameter count, so treat that figure as Decrypt's reporting.
The API specs, per xAI's release notes: a 500,000-token context window, text and image input, text output with no output limit, and four reasoning effort levels from low to xhigh. A Grok 4.7 Fast variant runs about twice as fast at twice the price. The model launched in Cursor and xAI's Grok Build tool, the xAI API, and model routers. GitHub added it to Copilot the same day for Pro, Pro+, Max, Business and Enterprise plans, billed at provider list pricing.
xAI's published benchmarks show clear improvement over Grok 4.6: 46.3% on CursorBench 4.0 (up from 40.4%), 71.0% on DeepSWE v1.1 (up from 65.2%), 64.0% on its EEBench electrical engineering test (up from 53.0%) and 56.7% on HealthBench Professional (up from 48.5%). On DeepSWE, xAI puts Grok 4.7 slightly ahead of Anthropic's Claude Fable 5.1 at 70.0%. These are all vendor-run numbers.
What independent testing found
Artificial Analysis scores Grok 4.7 at 46 on its Intelligence Index, two points above Grok 4.6. That is well above the median, but behind the leaders: Claude Fable 5.1 and GPT-6 each score 53, per The Decoder. On office-style work the gap is smaller. Artificial Analysis measured 1,657 Elo on AA-Briefcase, up 111 points, and 1,695 on GDPval-AA, up 90 points, against 1,678 and 1,735 for Claude Fable 5.1, per Decrypt.
Agentic coding shows the biggest gap. On Terminal-Bench 4.0, The Decoder reports Grok 4.7 at 26%, against 60% for OpenAI's GPT-6 Astra and 55% for Claude Fable 5.1. That puts it next to DeepSeek V4.1 Flash at 27%, a much cheaper model. For a release marketed first as a coding model, this is the weakest result.
The cost per task
This is where the flat price becomes misleading. Artificial Analysis reports Grok 4.7 at xhigh effort generated 240 million output tokens across its Intelligence Index run, against a median of 88 million across models. At high effort it generated 200 million. It also measured output speed at 44.6 tokens per second, which it describes as notably slow for the price tier.
Combined, that makes a cheap model more expensive to use. Artificial Analysis puts the cost at about $3.74 per Intelligence Index task at xhigh and $2.73 at high. VentureBeat notes that GPT-6 Astra at maximum effort used about 27,000 output tokens per task, about a third of Grok 4.7's 81,000. The point VentureBeat makes is simple: "A model charging less per token can still be more expensive on a finished workload."
xAI's own positioning says the model is "twice as fast, at half the price of comparable models." That can be true per token while being false per job. Buyers comparing models should look at cost per completed task at a matched reasoning setting, not the rate card.
Safety claims
xAI says Grok 4.7 ships with an entirely new safeguard stack and calls it the strongest model it has tested on refusals and jailbreak resistance. It reports 62.4% on LatchBio's biosafety benchmark and says only 3.3% of risky dual-use prompts got through on HackerBench v0.3, while legitimate security work was rarely blocked. xAI has also started giving select cybersecurity partners invite-only access to red-team functions. As with the capability numbers, these are xAI's own results. Independent safety evaluations have not been published yet.
Why it matters
Grok 4.7 is a solid improvement for xAI and a reasonable option for high-volume knowledge work where its Briefcase and GDPval scores sit close to the leaders. It does not close the gap on agentic coding, the category where the frontier labs are competing hardest this month.
The larger lesson is about how model pricing is now read. Reasoning models decide how long to think, and that decision can double a bill without any change to the posted price. Grok 4.7 is a clear example, but not the only one. Per-token pricing has become an incomplete measure, and cost per task at a stated effort level is the number that matters.
