Grok 4.7 arrived on September 21 with a pitch aimed squarely at anyone who watches their AI bill: a new, larger model at the same $2 input and $6 output price per million tokens as Grok 4.6, which had only shipped six weeks earlier.
The launch came with an unusually modest boast, too. Elon Musk wrote that it "places @SpaceXAI as third, after Anthropic & OpenAI, for agentic coding," and within two days both of those companies had shipped new models (the AI release calendar doesn't do quiet weeks anymore).
I went through xAI's announcement, its 29-page model card, the API docs and pricing tables, Artificial Analysis's independent evals, and the launch pages from Cursor, GitHub, AWS, and Oracle, so every number below traces back to its source.
If you're weighing a move to Grok 4.7 for coding or document work, or deciding whether this week's Claude and OpenAI releases are the better bet, the numbers below should settle it.
TL;DR:
Grok 4.7 is what SpaceXAI calls its most powerful model for coding and knowledge work, released on September 21, 2026, as the successor to Grok 4.6. If the name SpaceXAI is new to you, it's the name xAI now does business under, and the two are used interchangeably.
Developers call it with the model ID grok-4.7, and it has a 500,000-token context window, accepts text and images, and returns text. xAI's model docs give a knowledge cutoff of May 2026.
Cursor had a hand in this one. SpaceX acquired Cursor in August, and Cursor says it trained Grok 4.7 jointly with SpaceXAI on a new, larger base model.
xAI's model card adds that the model got supplemental training on anonymized Cursor workflow data to sharpen its coding.
The version numbers have been moving fast. Grok 4.5 launched on July 16, Grok 4.6 on August 12, and Grok 4.7 on September 21, so if you blinked this summer, you missed a model (possibly two).
One catch for chatbot users: Grok 4.7 isn't in grok.com, the Grok mobile apps, or Grok on X yet, and xAI says those consumer surfaces will get it at a later date. If you mainly use Grok as a chat assistant and you're weighing it against other AI tools like ChatGPT, you're still on Grok 4.6 for now.
The price and the context window stayed put, so the changes are all under the hood. Here's how the two compare on the points that affect your work and your bill:
It runs on a bigger base model. xAI says Grok 4.7 uses a new, larger base model than Grok 4.6. Its announcement and model card don't give a parameter count, though Musk has said on X that it has 2.1 trillion parameters. It still serves the model at the same price and speed as Grok 4.6.
It trained longer on harder, multi-hour tasks. The reinforcement learning run was weighted toward problems that take many hours to finish, and xAI says the result is a model that's better at checking its own work and managing long context.
Its effort levels are spaced further apart. Grok 4.6 already offered low, medium, high, and xhigh, but Cursor notes the gaps between them are wider on Grok 4.7, so harder settings spend more time thinking. Requests still default to high, and reasoning can't be switched off.
It's better at documents and presentations. xAI calls out document and slide creation as a specific improvement, which makes sense given that the model card also names it the default model in the Grok add-ins for Microsoft Word, PowerPoint, and Excel.
It has a new safety stack. xAI calls it the strongest model it has tested on refusals and jailbreak resistance, and says it let through only 3.3% of risky dual-use prompts on HackerBench. Select cybersecurity partners also get invite-only access to its red-team capabilities.
It thinks out loud a lot more. Artificial Analysis found it uses about 81K output tokens per Intelligence Index task, compared with 36K for Grok 4.6 at high effort. That's the change that shows up on your invoice, because you pay for every one of those extra tokens.
xAI's launch post compares Grok 4.7 at xhigh effort with Grok 4.6 at high effort. These are xAI's own numbers, so read them as the vendor's best case:
Terminal-Bench is the headline jump, nearly doubling from 20.3% to 37.6% in xAI's launch post. You may see 38.0% quoted too, because that's the figure in xAI's model card, which ran the test inside its Grok Build harness.
Legal work is its strongest showing against rivals. On the Harvey Legal Agent benchmark, xAI's table puts Grok 4.7 at 19.6% against 6.7% for Claude Fable 5.1 and 2.5% for GPT-5.6 Sol. Its electrical engineering score on EEBench (64.0%) also beats both of those models in xAI's table.
Artificial Analysis, which runs every model through the same test suite, scored Grok 4.7 at 46 on its Intelligence Index, two points above Grok 4.6, and said the result puts SpaceXAI among the top four AI labs.
Paired with Grok Build, it scored 56 on the Coding Agent Index, up from 47 with Grok 4.6, which ranks it fourth behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. That lines up neatly with Musk's "third" claim, since two of the three models ahead of it come from Anthropic.
Artificial Analysis's write-up also turned up one improvement and two weak spots:
Grok 4.7 is billed per million tokens on the xAI API, and the official pricing table has two tiers based on how long your prompt is:
The long-context rate applies to the whole request, so a 210K-token prompt pays double on all 210K tokens, including the first 200K.
Five other charges and limits sit outside that table:
For the consumer app, SuperGrok plans on grok.com run $10 (Lite), $30, $100 (Plus), and $300 (Heavy) a month, which is handy context if you're comparing Grok with other AI platforms. Just remember those plans still chat with Grok 4.6 until 4.7 reaches the app.
Grok 4.7's per-token price is low, but it spends a lot more tokens getting to an answer. Artificial Analysis measured about 81K output tokens per task, compared with 36K for Grok 4.6 at high effort, which works out to roughly 125% more output for the same job.
On Artificial Analysis's Intelligence Index, that makes Grok 4.7 cost $3.74 per task at xhigh and $2.73 at high, so the effort level you pick moves your bill about as much as the price sheet does.
xAI's own CursorBench chart tells a friendlier story for long coding jobs. There, Grok 4.7 at xhigh scored 46.3% at $6.01 per task, while Claude Fable 5.1 at max scored 51.8% at $17.28, so you pay about a third of the price for most of the score.
My suggestion is to start at high or medium effort and only reach for xhigh on the jobs that fail at lower settings. At low effort, xAI's chart shows Grok 4.7 hitting 33.1% on CursorBench for $1.58 a task, which is plenty for routine edits.
Grok 4.7 went out to developers and coding tools first, and cloud platforms followed within a week. Here's where you can use it as of September 29, 2026:
GitHub Copilot is still rolling out. GitHub's changelog says Grok 4.7 is available in Copilot on paid plans from Pro up to Enterprise, billed at xAI's list price, with a gradual rollout. If it's missing from your picker in VS Code, JetBrains, or the Copilot CLI, give it a few days.
The cloud platforms each use their own model ID. Amazon Bedrock lists us.xai.grok-4.7 and global.xai.grok-4.7, while Oracle's release notes list xai.grok-4.7, so check the exact string before you swap it into your config.
On the xAI API, Grok 4.7 works with both the Responses API and the Chat Completions API. You set the model to grok-4.7, pick an effort level with reasoning_effort (it defaults to high), and you can use function calling, structured outputs, web search, X search, and code execution.
The rate limits are generous, at 150 requests per second and 50 million tokens per minute, and the model runs in three US regions. If you're coming from Grok 4.6, the price is identical, so the main thing to retest is how many tokens your prompts now burn.
Coding is where xAI aimed this release. Grok 4.7 is the default model in Grok Build, SpaceXAI's terminal-based coding agent, which you can run as an interactive terminal app, headlessly in scripts, or inside other tools through the Agent Client Protocol.
Grok Build has a free tier, which makes it an easy way to kick the tires on Grok 4.7 without paying for API tokens (the Fast variant is excluded from that tier). Artificial Analysis also scored the model and the agent together, which is where that fourth-place Coding Agent Index result comes from.
Grok Build's closest rival is Anthropic's Claude Code, and the two make opposite trades between price and top-end coding scores, so it's worth comparing Grok Build and Claude Code before you commit to either agent.
Cursor is its other big home. Grok 4.7 is on every Cursor plan, per xAI, with a 256K standard window and a 500K long-context option, and Fast is the default speed on Pro and higher plans. That 256K figure is Cursor's own window, so don't mix it up with the API's 200K pricing threshold.
When you're sizing Grok 4.7 up against the other AI coding agents you already use, the simplest test is to run the same few real tasks in Cursor or Copilot with Grok 4.7 and then with your current model, comparing both the result and the token count.
{{templates}}
The rival lineup changed within two days of Grok 4.7's launch, when Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol, which GPT-6.1 Sol replaced a week later.
Here's how Grok 4.7 compares with the current flagships, using each vendor's own figures as of September 2026:
Grok 4.7 has the lowest output price in the group, and it ties GPT-6.1 Sol on input. Against the top-end models, it's a fifth of the input price of Fable 5.1 and GPT-6 Astra, and less than an eighth of their output price.
Claude Opus 5.5 leads on terminal work by a wide margin. Anthropic reports 66.4% on Terminal-Bench 4.0 for Opus 5.5, against Grok 4.7's 37.6%, and 57.8% on CursorBench 4.0 against 46.3%. It costs twice as much on input and more than three times as much on output.
GPT-6.1 Sol is the closest price match. OpenAI released it on September 29 to replace GPT-6 Sol, at the same $2 input and $10 output, half the price of GPT-5.6 Sol. OpenAI hasn't published a Terminal-Bench 4.0 score for it, so test both on your own tasks for now.
Shopping across model families anyway? Read up on Claude alternatives before you commit to any one of them, since the lineup keeps shuffling month to month.
Grok 4.7 is a strong value pick, but it has a few weak spots you should know before you switch:
{{cta}}
Grok 4.7 is one of the cheapest ways to get near-frontier coding and knowledge work from a major lab right now, as long as you watch how many tokens it spends.
Choose Grok 4.7 if you:
Skip it for now if you:
If your team runs everyday work through an AI teammate like Lindy, which lives in Slack and is model agnostic, check which models your workspace lists before you plan anything around Grok 4.7.
For everyone else, the cheapest experiment is also the most useful one: switch your model picker to Grok 4.7 for a week on the tasks you already run, keep an eye on the token counts, and you'll know whether the low sticker price holds up for your work.
Grok 4.7 is free to try in Grok Build's free tier. The xAI API is paid at $2 per million input tokens and $6 per million output tokens, and the Grok chat app doesn't have Grok 4.7 yet.
Grok 4.7 is the current Grok model, released on September 21, 2026. It's live on the xAI API, Grok Build, Cursor, and GitHub Copilot, while grok.com and the Grok apps still run Grok 4.6 until 4.7 reaches consumer surfaces.
Grok 4.6 and Grok 4.5 are both still available on the xAI API, at the same $2 input and $6 output price as Grok 4.7. Grok 4.6 launched on August 12, 2026, and Grok 4.5 on July 16, 2026, and xAI hasn't announced a retirement date for either.
The Grok 4.7 context window is 500,000 tokens on the xAI API. Prompts of 200K tokens or more are billed at double the standard rate, and Cursor uses its own 256K standard window with a 500K long-context option.
Grok 4.7 is cheaper than Claude, but Claude's newest models score higher on key coding benchmarks. Claude Opus 5.5 reports 66.4% on Terminal-Bench 4.0 against Grok 4.7's 37.6%, while Grok 4.7 costs half as much on input. Weighing OpenAI too? See how Claude compares with ChatGPT.
