AI Explained

xAI's Grok 4.7: Cheap Sticker Price, Pricier Real Bill

Arpit TripathiArpit TripathiLinkedIn·September 23, 2026·12 min read

Grok 4.7 shipped with cheap $2/$6 pricing but a pricier real bill than GPT-6 Astra per task. Benchmarks, pricing, and who should switch.

xAI missed its own Grok 4.7 deadline at least five times before finally shipping it on September 21, 2026. The pricing holds steady at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, while Terminal-Bench 4.0, a benchmark built around complex, multi-step engineering work, jumps from 20.3% to roughly 38%, the model's single largest reported gain over its own predecessor. That's a much cheaper option than the current frontier leaders on paper, but Artificial Analysis found Grok 4.7 actually costs more per completed task than OpenAI's GPT-6 Astra once real-world verbosity is factored in, and it does not overtake Claude's Fable 5.1 or GPT-6 Astra on the benchmarks that measure real professional and coding work.

The slippage started earlier than most coverage credits it, and it followed a consistent pattern: Musk kept naming a new date, then missing it.

  • July 25: Musk hints at an August 22 release target.
  • August 12, the day Grok 4.6 ships: revised to "three to four weeks" out.
  • Late August: that window slides again, to early September.
  • September 2: Musk gives "10 days."
  • September 11: "needs a few more days to cook," ten days before the model actually ships.
  • September 21: Grok 4.7 finally ships.

That pattern shapes how to read the launch. A model that slips this many times before arriving invites two different reactions once it does: either the extra months bought something worth the wait, or the company kept missing its own targets until the work was finally done. Grok 4.7's numbers support a version closer to the second reading, solid and incremental rather than the leap the repeated delays might have implied.

Grok 4.7: What Actually Shipped on September 21

xAI describes Grok 4.7's training as starting from "a new, larger base than 4.6, then a longer reinforcement learning run on a harder mix weighted toward problems that take many hours," aimed at improving self-verification and long-context management rather than simply scaling parameters. Independent coverage, though not xAI's own launch page in its exact words, reports the training run also drew on SpaceX operational data, Starlink satellite telemetry, manufacturing records, and engineering fault logs, intended to strengthen the model's reasoning about hardware and physical systems rather than relying only on internet text. xAI has not confirmed that detail in its own launch post, so treat it as well-corroborated secondary reporting rather than an official spec.

The SpaceX Data Bet

Grok 4.7's delay ties directly to that training decision, according to Winzheng's reporting: the outlet cites the slip to "additional training on the model, injecting large volumes of SpaceX engineering data" as the stated reason for the slide from Musk's original three-to-four-week estimate. The specific categories named across independent coverage are Starlink satellite telemetry, rocket development records and test logs, and internal engineering documents, not a single dataset but a deliberate mix of operational data from a hardware company rather than another pass over internet text. The bet is straightforward to state and hard to verify from outside xAI: that a model trained partly on how real engineering systems fail and get fixed reasons better about hardware and physical-systems problems than one trained only on text scraped from the web. Grok 4.7's benchmark gains land mostly in coding and terminal-based tasks rather than a named physical-systems benchmark, so whether the SpaceX data bet paid off in a way outsiders can measure remains open.

The parameter count that appeared across nearly every outlet covering the launch, 2.1 trillion, up 40% from Grok 4.6's reported 1.5 trillion, does not appear on xAI's own model card or launch post. It traces back to Musk, who was already citing the figure in posts as early as late July and again on September 2, weeks before launch, and independent trackers including llm-stats.com note they could not confirm the figure directly from xAI's technical documentation. Treat 2.1 trillion as Musk's stated number for now, not a confirmed xAI spec, the same caution worth applying to any parameter count a lab has not put in its own technical writeup.

Grok 4.7 launched simultaneously across five surfaces:

  • The Grok app, with no waitlist
  • Cursor, the code editor, as a named launch partner
  • Grok Build, xAI's own coding environment
  • The xAI API, priced the same as chat access
  • Selected cybersecurity partners, given early red-team access ahead of the public launch

Grok 4.7 Pricing: Holds Steady Until 200,000 Tokens

Grok 4.7's headline price, $2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens, is identical to Grok 4.6's, and cached input costs $0.50 per million tokens under that same threshold. One detail sits past that line, one most users will not hit but should know about: once a single prompt crosses 200,000 tokens, the entire request bills at a higher tier, not just the tokens past the threshold. Above 200,000 tokens, input runs $4 per million, cached input $1 per million, and output $12 per million.

Context window sits at 500,000 tokens, the same class as Grok 4.6 and half of the 1 million tokens Fable 5.1 carries. Pretraining cutoff is reported slightly differently across sources: llm-stats.com puts it at June 2026 with supplemental training through August, while iweaver.ai reports May 2026. The two estimates are close enough to call the cutoff roughly mid-2026 with confidence, but not close enough to treat a single exact month as settled fact.

Grok 4.7 Benchmarks: Terminal-Bench Is the Real Story

Every benchmark xAI published for Grok 4.7 improved over Grok 4.6, and the size of the gain varies a lot by task. CursorBench 4.0 climbed from 40.4% to 46.3%. DeepSWE v1.1 reached 71.0% at high effort, up from 65.2%. EEBench gained ground too, from 53.0% to roughly 64 to 66%. The Harvey Legal Agent Benchmark moved from 15.8% to 19.6%, and HealthBench Professional landed at 56.7%. SWE-Marathon v1.1, a longer-horizon coding test, jumped from 31.9% to 46.0%, a gain second only to Terminal-Bench in size and further evidence that the longer reinforcement-learning run specifically targeted tasks that take longer to finish rather than short single-turn prompts. The Terminal-Bench 4.0 jump dwarfs all of those: 20.3% to roughly 38%, a gain of about 17 to 18 points, xAI's own largest reported improvement and the clearest evidence that the longer reinforcement-learning run weighted toward multi-hour tasks changed how the model handles extended, multi-step work inside a real terminal.

Grok 4.7's benchmark gains do not add up to leading the field, and the picture against the other two frontier labs is genuinely mixed rather than one-sided. On GDPval-AA, Artificial Analysis's professional-work index, Grok 4.7's xhigh setting scores 1,695, behind Fable 5.1's 1,735 but ahead of GPT-6 Astra's 1,542. AA-Briefcase, another Artificial Analysis measure, puts Grok 4.7 at 1,657 against Fable 5.1's 1,678. On CursorBench 4.0, Grok 4.7's 46.3% beats the GPT-5.6 Sol coding-tier variant's 41.7% but trails Fable 5.1's 51.8%. On Terminal-Bench 4.0 specifically, the benchmark where Grok 4.7 made its biggest jump over its own predecessor, GPT-6 Astra still leads outright at 57.9%, ahead of both Fable 5.1's 55.8% and Grok 4.7's 38.0%. Zoom out to Artificial Analysis's broader Intelligence Index, a composite score, and Grok 4.7 sits at 46 against 53 for both GPT-6 Astra and Fable 5.1. Coverage of the launch has settled on describing Grok 4.7 as mid-range on coding work: it wins some individual matchups against Astra and Sol, loses others, and does not lead the field outright on any of the widely cited composite indexes.

The price comparison gets more complicated once actual task cost, not just the per-token rate, enters the picture. Artificial Analysis, which runs its own fixed test set against each model rather than relying on the headline price sheet, found Grok 4.7 costs $2.73 per completed task on its Intelligence Index run, against $0.82 for GPT-6 Astra's cheapest setting, more than three times as much despite Grok's per-token rate running roughly a fifth of Astra's. Grok 4.7's higher real-world cost per task traces back to verbosity: it needs more processing steps and more output tokens to reach the same result on that test set, so a lower sticker price does not automatically mean a cheaper bill once a model is actually working through a task.

Grok 4.7 Reception: Cheaper and Improved, Not the Leap

Coverage of the launch has landed on a consistent frame: Grok 4.7 arrives late to the current AI frontier cycle, competent and meaningfully cheaper than the leaders, but not the model that resets where the frontier sits. Grok 4.7's numbers line up with that reading: a real 17 to 18 point jump on the one benchmark that moved the most, unchanged pricing that stays well below Fable 5.1 on a per-token basis, and a consistent second-place finish against Fable 5.1 on the professional-work and coding indexes cited most often. Grok 4.6 to Grok 4.7 is the comparison the numbers actually support cleanly. Grok 4.7 against the rest of the current frontier is a closer call than the headline pricing alone suggests, once real task cost gets factored in.

What You're ComparingGrok 4.7Grok 4.6Claude Fable 5.1
Price per million tokens (input / output, under 200K)$2 / $6$2 / $6$10 / $50
Context window500K tokens500K tokens (same class)1M tokens
Terminal-Bench 4.0~37.6-38.0%20.3%55.8%
GDPval-AA (professional-work index)1,695 (xhigh)not published by xAI1,735
Best fit forCost-sensitive coding and hardware-adjacent tasks where per-token price matters mostExisting Grok 4.6 deployments with no urgent reason to moveHighest-ceiling reasoning and long-horizon professional work

Who Should Actually Consider Grok 4.7

  • Existing Grok 4.6 users doing coding or terminal-heavy agent work: the Terminal-Bench and SWE-Marathon gains are large enough, at unchanged pricing, that this is close to a straightforward upgrade with little reason to wait.
  • Teams already paying for Cursor who want a materially cheaper model in the rotation: Grok 4.7 shipped in Cursor on day one, and its per-token price sits well below Fable 5.1 and GPT-6 Astra even if its cost per completed task narrows that gap somewhat.
  • Anyone building tools that reason about hardware, manufacturing, or physical-systems failures: the SpaceX-derived training data is the one differentiator no other frontier model has claimed, even though its real-world effect is not independently measurable yet.
  • Teams that need the highest ceiling on professional and legal work and can absorb the cost: Fable 5.1 still leads Grok 4.7 on GDPval-AA and AA-Briefcase, though Grok 4.7 actually leads decisively on the Harvey Legal Agent Benchmark (19.6% versus Fable 5.1's 6.7%), so it is the general professional-work gap, not legal-agent performance, that is the reason to pay more rather than switch.
  • Anyone comparing models purely on the per-token price sheet: Artificial Analysis's cost-per-task numbers are a reminder to test on an actual workload before assuming the cheaper rate card produces the cheaper bill.

Where MemX Fits: Memory That Does Not Reset With Every Model Swap

Every one of those numbers still leaves a problem that shows up regardless of which model ships next: a Grok subscription, a Claude subscription, and a ChatGPT subscription each start every new conversation blind to what happened in the other two. Grok 4.7's SpaceX-flavored engineering data and its Terminal-Bench gains make it a reasonable pick for hardware-adjacent or long-running coding tasks priced well below Fable 5.1, while Fable 5.1 still holds the edge on the professional-work benchmarks that matter for other kinds of work, and GPT-6 Astra's own cost-per-task numbers make it worth a look for anyone optimizing purely on price. Someone routing tasks between three models by strength and price, rather than living inside one vendor's app, is exactly who ends up re-explaining the same project context, the client history, the decisions already made, the reasons a prior approach got ruled out, to a second or third model every week. MemX (memx.app) is built for that specific gap: a memory layer that sits outside any single provider's app, kept private by architecture, so context built up while working with one model does not have to be retyped from scratch when the next task calls for a different one.

Frequently Asked Questions
01When did xAI release Grok 4.7 and why was it delayed so many times?

xAI released Grok 4.7 on Monday, September 21, 2026. Elon Musk pushed the release date back at least five times since late July: an August 22 target on July 25, "three to four weeks" out on August 12, a slip to early September, "10 days" on September 2, and "needs a few more days to cook" on September 11, ten days before it actually shipped.

02How much does Grok 4.7 cost?

$2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens, the same as Grok 4.6. Cached input costs $0.50 per million under that threshold. Once a prompt crosses 200,000 tokens, the entire request bills at a higher tier: $4 per million input, $1 per million cached input, and $12 per million output.

03Does Grok 4.7 really have 2.1 trillion parameters?

Grok 4.7's parameter count is widely reported as 2.1 trillion, but that figure traces back to posts Elon Musk made weeks before launch, as early as late July and again on September 2, rather than xAI's own model card or technical documentation, and independent trackers including llm-stats.com say they could not confirm it directly. Treat 2.1 trillion, up 40% from Grok 4.6's reported 1.5 trillion, as Musk's stated figure rather than a confirmed xAI spec.

04Is Grok 4.7 better than Claude Fable 5.1 or GPT-6 Astra?

Not outright. Grok 4.7 improves on every benchmark xAI published against its own predecessor, most sharply on Terminal-Bench 4.0, but it trails Fable 5.1 on GDPval-AA (1,695 versus 1,735), AA-Briefcase (1,657 versus 1,678), and CursorBench 4.0 (46.3% versus 51.8%). Its clearest advantage is price: its per-token rate runs well below Fable 5.1's and GPT-6 Astra's, though Artificial Analysis found its actual cost per completed task runs higher than GPT-6 Astra's cheapest setting once verbosity is factored in.

05Where can I use Grok 4.7?

It shipped live on September 21 with no waitlist, in the Grok app, in Cursor, in Grok Build, and through the xAI API. Selected cybersecurity partners also received early red-team access ahead of the public launch.

Read Next

Or try MemX to access 40+ AI models in one place — including Claude Sonnet 4.6 and GPT-5.4 — and get your questions answered today.

Was this article helpful?

Found this useful? Share it with someone who needs it.

Free · iOS, Android & WhatsApp

Stop losing what you save.
Let MemX remember it for you.

Every screenshot, photo, PDF and voice note — captured, encrypted, and instantly searchable. Ask in plain English, get the answer in seconds.

  • Reads text inside images and handwriting
  • Private and encrypted by default
  • Free to start, no credit card

Takes under a minute to set up. Your data stays yours.

Arpit Tripathi
Written by
Arpit TripathiLinkedIn

Founder of MemX. Ex-Google Staff Tech Lead Manager, ex-AWS Senior SDE (Elastic Block Store). Writes about practical AI on the MemX blog.

Keep reading

More guides for AI-powered students.