1. Home
  2. Blog
  3. AI Tools

Grok 4.7: Pricing, Benchmarks, and What Changed From 4.6

Lindy Drope
Lindy Drope
Founding GTM at Lindy
Lindy leads GTM at Lindy and is the team’s most prolific automation builder. She publishes weekly educational videos and articles on building AI assistants – And yes, she’s a real person!
Lindy Drope
Written by
Lindy Drope
Marvin Aziz
Marvin Aziz
Growth Engineer
Marvin is a Growth Engineer at Lindy focused on AI agents, automation, and product-led growth.
Marvin Aziz
Reviewed by
Marvin Aziz
Last Updated:
October 1, 2026
Expert Verified

Grok 4.7 arrived on September 21 with a pitch aimed squarely at anyone who watches their AI bill: a new, larger model at the same $2 input and $6 output price per million tokens as Grok 4.6, which had only shipped six weeks earlier.

The launch came with an unusually modest boast, too. Elon Musk wrote that it "places @SpaceXAI as third, after Anthropic & OpenAI, for agentic coding," and within two days both of those companies had shipped new models (the AI release calendar doesn't do quiet weeks anymore).

I went through xAI's announcement, its 29-page model card, the API docs and pricing tables, Artificial Analysis's independent evals, and the launch pages from Cursor, GitHub, AWS, and Oracle, so every number below traces back to its source.

If you're weighing a move to Grok 4.7 for coding or document work, or deciding whether this week's Claude and OpenAI releases are the better bet, the numbers below should settle it.

TL;DR:

  • Released: September 21, 2026, as the successor to Grok 4.6.
  • What it is: SpaceXAI's flagship model for coding and knowledge work, built on a new, larger base model.
  • Price: $2 per million input tokens and $6 per million output tokens, doubling once a prompt reaches 200K tokens.
  • Context window: 500,000 tokens, with text and image input.
  • Where to use it: the xAI API, Grok Build, Cursor, GitHub Copilot, OpenRouter, Vercel, Cloudflare, Oracle Cloud, and Amazon Bedrock. It isn't in the Grok app yet.
  • The catch: it uses about twice as many output tokens per task as Grok 4.6, so cheap tokens don't always mean a cheap job.

What is Grok 4.7?

Grok 4.7 is what SpaceXAI calls its most powerful model for coding and knowledge work, released on September 21, 2026, as the successor to Grok 4.6. If the name SpaceXAI is new to you, it's the name xAI now does business under, and the two are used interchangeably.

Developers call it with the model ID grok-4.7, and it has a 500,000-token context window, accepts text and images, and returns text. xAI's model docs give a knowledge cutoff of May 2026.

Cursor had a hand in this one. SpaceX acquired Cursor in August, and Cursor says it trained Grok 4.7 jointly with SpaceXAI on a new, larger base model.

xAI's model card adds that the model got supplemental training on anonymized Cursor workflow data to sharpen its coding.

The version numbers have been moving fast. Grok 4.5 launched on July 16, Grok 4.6 on August 12, and Grok 4.7 on September 21, so if you blinked this summer, you missed a model (possibly two).

One catch for chatbot users: Grok 4.7 isn't in grok.com, the Grok mobile apps, or Grok on X yet, and xAI says those consumer surfaces will get it at a later date. If you mainly use Grok as a chat assistant and you're weighing it against other AI tools like ChatGPT, you're still on Grok 4.6 for now.

What's new in Grok 4.7 vs Grok 4.6

The price and the context window stayed put, so the changes are all under the hood. Here's how the two compare on the points that affect your work and your bill:

🔍 Spec ⏮️ Grok 4.6 ⏭️ Grok 4.7
Released August 12, 2026 September 21, 2026
Base model Previous base New, larger base
Input / output price $2 / $6 per 1M $2 / $6 per 1M
Context window 500K tokens 500K tokens
Terminal-Bench 4.0 20.3% 37.6%
Output tokens per AA task About 36K About 81K

It runs on a bigger base model. xAI says Grok 4.7 uses a new, larger base model than Grok 4.6. Its announcement and model card don't give a parameter count, though Musk has said on X that it has 2.1 trillion parameters. It still serves the model at the same price and speed as Grok 4.6.

It trained longer on harder, multi-hour tasks. The reinforcement learning run was weighted toward problems that take many hours to finish, and xAI says the result is a model that's better at checking its own work and managing long context.

Its effort levels are spaced further apart. Grok 4.6 already offered low, medium, high, and xhigh, but Cursor notes the gaps between them are wider on Grok 4.7, so harder settings spend more time thinking. Requests still default to high, and reasoning can't be switched off.

It's better at documents and presentations. xAI calls out document and slide creation as a specific improvement, which makes sense given that the model card also names it the default model in the Grok add-ins for Microsoft Word, PowerPoint, and Excel.

It has a new safety stack. xAI calls it the strongest model it has tested on refusals and jailbreak resistance, and says it let through only 3.3% of risky dual-use prompts on HackerBench. Select cybersecurity partners also get invite-only access to its red-team capabilities.

It thinks out loud a lot more. Artificial Analysis found it uses about 81K output tokens per Intelligence Index task, compared with 36K for Grok 4.6 at high effort. That's the change that shows up on your invoice, because you pay for every one of those extra tokens.

Grok 4.7 benchmarks

xAI's launch post compares Grok 4.7 at xhigh effort with Grok 4.6 at high effort. These are xAI's own numbers, so read them as the vendor's best case:

📊 Benchmark ⏮️ Grok 4.6 ⏭️ Grok 4.7
CursorBench 4.0 40.4% 46.3%
DeepSWE v1.1 65.2% 71.0% (high)
Terminal-Bench 4.0 20.3% 37.6%
EEBench 53.0% 64.0%
AA Briefcase v1.1 1,546 Elo 1,657 Elo
GDPval 1,605 Elo 1,695 Elo
Harvey Legal Agent 15.8% 19.6%
HealthBench Professional 48.5% 56.7%

Terminal-Bench is the headline jump, nearly doubling from 20.3% to 37.6% in xAI's launch post. You may see 38.0% quoted too, because that's the figure in xAI's model card, which ran the test inside its Grok Build harness.

Legal work is its strongest showing against rivals. On the Harvey Legal Agent benchmark, xAI's table puts Grok 4.7 at 19.6% against 6.7% for Claude Fable 5.1 and 2.5% for GPT-5.6 Sol. Its electrical engineering score on EEBench (64.0%) also beats both of those models in xAI's table.

What independent testing shows

Artificial Analysis, which runs every model through the same test suite, scored Grok 4.7 at 46 on its Intelligence Index, two points above Grok 4.6, and said the result puts SpaceXAI among the top four AI labs.

Paired with Grok Build, it scored 56 on the Coding Agent Index, up from 47 with Grok 4.6, which ranks it fourth behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. That lines up neatly with Musk's "third" claim, since two of the three models ahead of it come from Anthropic.

Artificial Analysis's write-up also turned up one improvement and two weak spots:

  • Fewer hallucinations: its hallucination rate on AA-Omniscience fell to 29%, from 34% for Grok 4.6.
  • Some regressions: it slipped 3.7 points on AA-LCR, a long-context reasoning test, and 1.1 points on AutomationBench-AA.
  • Lower Terminal-Bench on a neutral harness: Artificial Analysis's own standardized run puts it at about 26%, well below xAI's 37.6%, which is a good reminder that harness choice moves these scores a lot.

Grok 4.7 pricing

Grok 4.7 is billed per million tokens on the xAI API, and the official pricing table has two tiers based on how long your prompt is:

💵 Price per 1M tokens 📄 Prompt under 200K 📚 Prompt 200K or more
Input $2.00 $4.00
Cached input $0.50 $1.00
Output $6.00 $12.00

The long-context rate applies to the whole request, so a 210K-token prompt pays double on all 210K tokens, including the first 200K.

Five other charges and limits sit outside that table:

  • Grok 4.7 Fast: the same model on faster infrastructure, at $4 input, $1 cached, and $12 output per million tokens. It's only available in Cursor and Grok Build.
  • US regional endpoint: 10% more, at $2.20 input and $6.60 output, for workloads that need to run in the US.
  • Priority processing: billed at twice the standard rate.
  • Web search tool: $5 per 1,000 calls.
  • Batch API: not supported for Grok 4.7, so there's no batch discount for jobs that can wait.

For the consumer app, SuperGrok plans on grok.com run $10 (Lite), $30, $100 (Plus), and $300 (Heavy) a month, which is handy context if you're comparing Grok with other AI platforms. Just remember those plans still chat with Grok 4.6 until 4.7 reaches the app.

Why cheap tokens can still mean a pricier task

Grok 4.7's per-token price is low, but it spends a lot more tokens getting to an answer. Artificial Analysis measured about 81K output tokens per task, compared with 36K for Grok 4.6 at high effort, which works out to roughly 125% more output for the same job.

On Artificial Analysis's Intelligence Index, that makes Grok 4.7 cost $3.74 per task at xhigh and $2.73 at high, so the effort level you pick moves your bill about as much as the price sheet does.

xAI's own CursorBench chart tells a friendlier story for long coding jobs. There, Grok 4.7 at xhigh scored 46.3% at $6.01 per task, while Claude Fable 5.1 at max scored 51.8% at $17.28, so you pay about a third of the price for most of the score.

My suggestion is to start at high or medium effort and only reach for xhigh on the jobs that fail at lower settings. At low effort, xAI's chart shows Grok 4.7 hitting 33.1% on CursorBench for $1.58 a task, which is plenty for routine edits.

How to access Grok 4.7

Grok 4.7 went out to developers and coding tools first, and cloud platforms followed within a week. Here's where you can use it as of September 29, 2026:

🧭 Where 🛠️ How to get it ⚠️ Notes
xAI API Model ID grok-4.7 Paid, $2 / $6 per 1M
Grok Build Default model Free tier available
Cursor All plans Fast is default on Pro and up
GitHub Copilot Model picker Pro, Pro+, Max, Business, Enterprise
OpenRouter x-ai/grok-4.7 Same list price
Vercel AI Gateway spacexai/grok-4.7 Launch discount has ended
Cloudflare xai/grok-4.7 Zero data retention
Oracle Cloud xai.grok-4.7 Since September 25
Amazon Bedrock us.xai.grok-4.7 Since September 28
Grok app and X Not yet Coming "at a later date"

GitHub Copilot is still rolling out. GitHub's changelog says Grok 4.7 is available in Copilot on paid plans from Pro up to Enterprise, billed at xAI's list price, with a gradual rollout. If it's missing from your picker in VS Code, JetBrains, or the Copilot CLI, give it a few days.

The cloud platforms each use their own model ID. Amazon Bedrock lists us.xai.grok-4.7 and global.xai.grok-4.7, while Oracle's release notes list xai.grok-4.7, so check the exact string before you swap it into your config.

Using the Grok 4.7 API

On the xAI API, Grok 4.7 works with both the Responses API and the Chat Completions API. You set the model to grok-4.7, pick an effort level with reasoning_effort (it defaults to high), and you can use function calling, structured outputs, web search, X search, and code execution.

The rate limits are generous, at 150 requests per second and 50 million tokens per minute, and the model runs in three US regions. If you're coming from Grok 4.6, the price is identical, so the main thing to retest is how many tokens your prompts now burn.

Grok 4.7 for coding and Grok Build

Coding is where xAI aimed this release. Grok 4.7 is the default model in Grok Build, SpaceXAI's terminal-based coding agent, which you can run as an interactive terminal app, headlessly in scripts, or inside other tools through the Agent Client Protocol.

Grok Build has a free tier, which makes it an easy way to kick the tires on Grok 4.7 without paying for API tokens (the Fast variant is excluded from that tier). Artificial Analysis also scored the model and the agent together, which is where that fourth-place Coding Agent Index result comes from.

Grok Build's closest rival is Anthropic's Claude Code, and the two make opposite trades between price and top-end coding scores, so it's worth comparing Grok Build and Claude Code before you commit to either agent.

Cursor is its other big home. Grok 4.7 is on every Cursor plan, per xAI, with a 256K standard window and a 500K long-context option, and Fast is the default speed on Pro and higher plans. That 256K figure is Cursor's own window, so don't mix it up with the API's 200K pricing threshold.

When you're sizing Grok 4.7 up against the other AI coding agents you already use, the simplest test is to run the same few real tasks in Cursor or Copilot with Grok 4.7 and then with your current model, comparing both the result and the token count.

‍

{{templates}}

‍

Grok 4.7 vs Claude and OpenAI's latest models

The rival lineup changed within two days of Grok 4.7's launch, when Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol, which GPT-6.1 Sol replaced a week later. 

Here's how Grok 4.7 compares with the current flagships, using each vendor's own figures as of September 2026:

🤖 Model 💵 Input / output per 1M 📚 Context 📊 Terminal-Bench 4.0
Grok 4.7 $2 / $6 500K 37.6%
Claude Opus 5.5 $4 / $20 1M 66.4%
Claude Fable 5.1 $10 / $50 1M 55.8%
GPT-6 Astra $10 / $50 1.05M 57.9%
GPT-6.1 Sol $2 / $10 1.05M Not published

Grok 4.7 has the lowest output price in the group, and it ties GPT-6.1 Sol on input. Against the top-end models, it's a fifth of the input price of Fable 5.1 and GPT-6 Astra, and less than an eighth of their output price.

Claude Opus 5.5 leads on terminal work by a wide margin. Anthropic reports 66.4% on Terminal-Bench 4.0 for Opus 5.5, against Grok 4.7's 37.6%, and 57.8% on CursorBench 4.0 against 46.3%. It costs twice as much on input and more than three times as much on output.

GPT-6.1 Sol is the closest price match. OpenAI released it on September 29 to replace GPT-6 Sol, at the same $2 input and $10 output, half the price of GPT-5.6 Sol. OpenAI hasn't published a Terminal-Bench 4.0 score for it, so test both on your own tasks for now.

Shopping across model families anyway? Read up on Claude alternatives before you commit to any one of them, since the lineup keeps shuffling month to month.

Where Grok 4.7 falls short

Grok 4.7 is a strong value pick, but it has a few weak spots you should know before you switch:

  1. It's token-hungry. Doubling the output tokens per task eats into the low per-token price, and on some jobs a pricier model that thinks less can come out close.
  2. It trails the best models on long, hard coding. On FrontierSWE V2 in xAI's own model card, it scores 29.0% against 56.3% for Claude Fable 5.1, and it sits well behind Opus 5.5 on Terminal-Bench.
  3. The Fast variant isn't on the API. If you want the quicker version, you'll need Cursor or Grok Build, and it costs twice as much.
  4. Most benchmarks come from xAI. Independent runs on neutral harnesses can land lower, as Artificial Analysis's roughly 26% Terminal-Bench result shows.

‍

{{cta}}

‍

Is Grok 4.7 worth it? A quick decision guide

Grok 4.7 is one of the cheapest ways to get near-frontier coding and knowledge work from a major lab right now, as long as you watch how many tokens it spends.

Choose Grok 4.7 if you:

  • Run lots of coding or document work where cost per task is the number your team watches
  • Already work in Cursor, GitHub Copilot, or Grok Build and can test it with a model-picker change
  • Handle legal work, where xAI's Harvey benchmark puts it well ahead of Fable 5.1 and GPT-5.6 Sol
  • Want a 500K-token context window at $2 per million input tokens

Skip it for now if you:

  • Need the strongest results on long terminal and agentic coding runs, where Claude Opus 5.5 and Fable 5.1 still lead
  • Want Grok 4.7 in the Grok chat app, which hasn't gotten it yet
  • Need Batch API discounts or the Fast variant through the API

If your team runs everyday work through an AI teammate like Lindy, which lives in Slack and is model agnostic, check which models your workspace lists before you plan anything around Grok 4.7.

For everyone else, the cheapest experiment is also the most useful one: switch your model picker to Grok 4.7 for a week on the tasks you already run, keep an eye on the token counts, and you'll know whether the low sticker price holds up for your work.

FAQ

Is Grok 4.7 free?

Grok 4.7 is free to try in Grok Build's free tier. The xAI API is paid at $2 per million input tokens and $6 per million output tokens, and the Grok chat app doesn't have Grok 4.7 yet.

Which version of Grok is current?

Grok 4.7 is the current Grok model, released on September 21, 2026. It's live on the xAI API, Grok Build, Cursor, and GitHub Copilot, while grok.com and the Grok apps still run Grok 4.6 until 4.7 reaches consumer surfaces.

What happened to Grok 4.6 and Grok 4.5?

Grok 4.6 and Grok 4.5 are both still available on the xAI API, at the same $2 input and $6 output price as Grok 4.7. Grok 4.6 launched on August 12, 2026, and Grok 4.5 on July 16, 2026, and xAI hasn't announced a retirement date for either.

What is the Grok 4.7 context window?

The Grok 4.7 context window is 500,000 tokens on the xAI API. Prompts of 200K tokens or more are billed at double the standard rate, and Cursor uses its own 256K standard window with a 500K long-context option.

Is Grok better than Claude?

Grok 4.7 is cheaper than Claude, but Claude's newest models score higher on key coding benchmarks. Claude Opus 5.5 reports 66.4% on Terminal-Bench 4.0 against Grok 4.7's 37.6%, while Grok 4.7 costs half as much on input. Weighing OpenAI too? See how Claude compares with ChatGPT.

Save 2 Hours Every Day
Lindy is your ultimate AI assistant that manages inbox, meetings, and follow-ups—so you stay ahead of the chaos.
Try Lindy for Free
About the editorial team
Lindy Drope
Lindy Drope
Founding GTM at Lindy

Lindy leads GTM at Lindy and is the team’s most prolific automation builder. She publishes weekly educational videos and articles on building AI assistants – And yes, she’s a real person!

Marvin Aziz
Marvin Aziz
Growth Engineer

Marvin is a Growth Engineer at Lindy focused on AI agents, automation, and product-led growth.

Ready when you are.

Free to try. In your Slack in two minutes.

Try for free
7-day free trial • Cancel anytime