Google Is Paying SpaceX $920 Million a Month to Rent Servers. That’s the Grok 4.5 Story Nobody Told You.
When your competitors depend on your infrastructure to train their own models, the benchmark leaderboard stops being the most important number in the room.

On July 9, 2026, two things happened at once. OpenAI released GPT-5.6 Sol to the public. SpaceX AI released Grok 4.5. One week later, Moonshot AI launched Kimi K3, an open-weight model with 2.8 trillion parameters. In eight days, three frontier-class AI releases occurred, each of which can toe-to-toe with the best models globally.
The industry had never seen anything like it. And yet the story people will remember from those eight days is not the one about who scored highest on which benchmark.
It’s the story of a model that finished fourth, costs five times less than the models of its competitors, and quietly dismantled the assumption that rankings are what actually matter in AI right now.
A year ago, when Grok 4 launched, the headline wrote itself: it doubled the competition in a single release. A clean, undeniable breakout. This time, the story is more complex, more strategically layered, and in certain ways more unsettling for competitors than anything Grok 4 ever produced. Understanding why requires looking past the leaderboard.
Why Fourth Place Is the Most Dangerous Position Right Now
On the Artificial Analysis Intelligence Index, which measures AI models independently across reasoning, coding, knowledge work, and scientific tasks, Grok 4.5 lands in fourth place with a score of 54. The three models for it, Claude Fable 5, GPT-5.5, and Claude Opus 4.8, are all meaningfully stronger on the broadest general benchmarks. That is not spin. That is the actual ranking.
What the ranking does not capture is the cost. Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens. The three models ranked above it cost roughly five times as much. For many use cases, you are buying comparable output at one-fifth the price.
On specific coding evaluations, it does more than hold its own. Grok 4.5 achieved first place on SWE-Marathon, a benchmark for AI performance on real software engineering tasks in actual codebases, with a score of 29%, surpassing Opus 4.8 at 26% and Fable at 24%. On Terminal-Bench 2.1, an advanced terminal-use evaluation, it scores 83.3%, placing it within one point of both Fable 5 and GPT-5.5, the two highest scorers on that test. On broader evaluations like DeepSWE and SWE-Bench Pro, the models for it pull ahead. The picture is genuinely mixed, which is actually the honest version of the story.
What makes the cost argument concrete is token efficiency. Where Claude Opus 4.8 uses approximately 67,000 output tokens to complete a single SWE-Bench Pro task, Grok 4.5 completes the same task using around 15,000, a 4.2x reduction. A cheaper model that also consumes fewer tokens per job is not just a bargain on paper. It is a structural cost advantage that compounds at scale.
This maps directly onto a pattern that economists have tracked since the nineteenth century. William Stanley Jevons, an economist in England, observed in 1865: steam engines became more efficient with fuel, yet they consumed more coal overall. The machines became cheaper to run, so people ran more of them, and the market expanded. Economists call this the Jevons paradox. The same dynamic is playing out here. When powerful AI becomes significantly cheaper to operate, the result is not a smaller market for AI. The result is an explosion in usage, more applications built, more tasks delegated, and more revenue generated from the same infrastructure. At $2 per million tokens, Grok 4.5 is not just a product. It is a volume strategy.
The Infrastructure Nobody Else Can Replicate
The benchmark results are only one part of what happened here. The more consequential development took place six weeks earlier.
On June 16, 2026, SpaceX announced a $60 billion all-stock agreement to acquire Anysphere, the company behind Cursor, the AI-powered code editor used by approximately 4 million active developers worldwide. As Axios confirmed at the time, this is the largest acquisition of a venture-backed startup ever recorded.
Sixty billion dollars for a code editor sounds like an irrational number until you understand what the code editor produces. The cursor’s value is not the software. It is the data generated by 4 million developers writing actual code in real projects every day. Training data of that nature, live sessions from professional engineers solving actual problems, is qualitatively different from code scraped from public repositories. It is annotated in context and reflects how developers actually think through problems.
The logic follows an older playbook. When Facebook acquired Instagram in 2012 for $1 billion, the company had 13 employees and zero revenue. Everyone called it overpriced. Zuckerberg had not bought a product. He had bought a strategic position and a data source that would become a $100 billion advertising engine. SpaceX did not pay $60 billion for a code editor. It paid $60 billion for an intelligence pipeline that feeds directly into Grok.
Grok 4.5 was released three weeks after that acquisition closed. An engineer at SpaceX AI acknowledged publicly that the Cursor data was integrated through supplemental training rather than baked into the foundation model from the start of pre-training. That distinction matters. The model released on July 9 is already partly shaped by Cursor’s developer data, but not fully. Integration of this into pre-training is slated for the next model, which is when the substantial benefits are projected.
Which means what you saw from Grok 4.5 is not the ceiling of what this system can produce. It is the floor.
The hardware behind all of this sits in Memphis, Tennessee, at Colossus, SpaceX AI’s supercomputing complex. Colossus now spans three buildings, housing more than 250,000 GPUs, including 30,000 of NVIDIA’s latest GB200 units, with a total power capacity of 2 gigawatts. That is the equivalent of two nuclear reactors, enough electricity to power a city of a million people. The full output of that infrastructure goes toward one objective: training the model that feeds the coding tool that generates the data that trains the next model.
No other AI laboratory currently operates this loop. OpenAI does not own a code editor. Anthropic does not own a supercomputer at this scale. Google has pieces of the picture, but not under a single organizational roof. SpaceX AI assembled all three assets — a frontier model, a data-generating product with millions of professional users, and the compute to close the loop—in less than 12 months.


