A few months ago, if you’d told me Elon Musk’s AI model would sit at the same benchmark level as GPT and Claude, I would have laughed. Grok was the chatbot you got for free on X and forgot about. It was clever in a way that felt like a party trick and useless in a way that felt permanent. If you were doing proper work, you used Claude Code or ChatGPT.
Period.
And then, in about ten days in early August 2026, SpaceXAI dropped Grok Bot, released Grok 4.6, shipped Imagine 2.0, and closed the largest startup acquisition in venture capital history at $60 billion.
Sixty billion dollars for a code editor. That’s 50 percent more than what Musk paid for Twitter. And somehow, it might be the smartest money he’s spent.
The White Elephant That Became a Weapon
To understand why that price is actually rational, you need to understand the problem it solved.
When SpaceX absorbed xAI back in February 2026, it inherited Colossus 1, the Memphis supercomputer housing roughly 200,000 Nvidia GPUs across over 300 megawatts of power capacity, built in 122 days in what was genuinely one of the most impressive infrastructure deployments in recent tech history. The problem was that nobody wanted what it was running. Grok had buzz. Grok did not have users. GPU utilization sat at around 11 percent.
Hundreds of millions of dollars of hardware, doing almost nothing.
There’s a term for this in strategy: the white elephant. The expression comes from the kings of Siam. When a monarch wanted to ruin a rival without open conflict, he would gift them a white elephant. The animal was sacred and couldn’t be put to work, but it couldn’t be disposed of either. You had to feed it, house it, and watch it drain your treasury. Colossus 1 was SpaceX’s white elephant.
Too expensive to shut down, too underutilized to justify.
What Musk did next is probably his best strategic move since making rockets reusable.
Buying the Data, Not the Product
Cursor was founded in San Francisco in 2022 and quickly became the AI-native code editor that actually changed how developers worked.
It was the first to deeply integrate AI help into the development environment, before Claude Code, before Codex, before everyone else tried to replicate what Cursor had already proven. But Cursor had a gap of its own: mountains of developer interaction data, billions of lines of real-world coding conversations, and no compute infrastructure to train its own frontier models. It was renting from Anthropic and OpenAI, paying to use their models, essentially subsidizing its competitors’ data advantage.
SpaceX had the inverse problem.
With a cathedral of GPUs, there is no data worth training on.
The partnership started in April 2026 as a technology deal, with SpaceX holding an acquisition option at $60 billion or a standalone partnership at $10 billion. Cursor’s team began training jointly with SpaceX immediately, and the first output was Grok 4.5, which included what Cursor described as “trillions of tokens of Cursor data.” It was the first Grok model that wasn’t only a chatbot. It could actually reason about code.
On August 14, the acquisition formally closed.
All-stock deal.
Anysphere shareholders received 389,289,254 SpaceX Class A shares. Cursor’s team has been redistributed into SpaceXAI’s internal structures. The standalone Cursor editor remains active for now, but the brand will progressively fold into the Grok ecosystem. The co-founders are billionaires. And SpaceX’s white elephant just got the data it needed to justify its existence.
Grok 4.6 Is the First Proof the Loop Works
On August 12, two days before the acquisition closed, SpaceXAI released Grok 4.6. It is not a new foundation model. It is Grok 4.5 with an extended post-training run using new engineering data, much of it sourced directly from Cursor’s interaction logs. Reinforcement learning inside agentic environments. No new architecture or parameter count published.

But the results are impossible to wave away. On the Artificial Analysis Intelligence Index, which aggregates nine benchmarks into a single composite, Grok 4.6 scores 61. That ties GPT-5.6 Sol. Claude Fable 5 sits at 62. Claude Opus 5 at 63. This is the first time any Grok model has played in this range. Five months ago, Grok was an afterthought in serious AI conversations. Today it sits one or two points behind the best models on the planet.
On knowledge work specifically, the picture is even more striking. On the GDPVal benchmark measuring performance on real-world professional tasks, Grok 4.6 posts the highest published score. On HarveyLAB, which evaluates legal reasoning, Grok 4.6 scores 15.8 percent compared to 2.5 percent for GPT-5.6 Sol and 11.3 percent for Fable 5. Those are not marginal differences. That is a model that has ingested a specific type of high-quality professional data, and it shows.
Now, benchmarks never tell the entire story, and I’ve been writing about this long enough to know that a model that scores well on a chart still has to survive the adoption test. People don’t switch tools because a number went up. They switch when the experience gets better. But Grok 4.6 has something else going for it.
The Price Argument Is Brutal
Claude Opus 5 charges $5 per million input tokens and $25 per million output tokens. GPT-5.6 Sol charges $30 per million on the output side. Grok 4.6 comes in at $6 per million input and $18 per million output. That is at least 60 percent cheaper per token than its closest equivalents at the frontier level, and the per-task cost measured independently by Artificial Analysis.
While it has risen from 36 cents to 84 cents, reflecting the quality jump from 4.5, it remains remarkably competitive for a model that now sits in the same benchmark range as models costing substantially more to operate.
The strategy SpaceXAI is running is not subtle. The best quality-to-price ratio is at the frontier. This is Musk’s playbook transplanted directly from SpaceX’s approach to launch costs: don’t be marginally cheaper, be structurally cheaper, and force the market to recalculate.
Grok Bot, Imagine 2.0, and the Distribution Problem
A well-placed model on a benchmark chart means nothing without distribution. Anthropic understood this before anyone, which is why Claude Code exists. Developers use Claude Code, which generates interaction data that feeds back into model training; the revenue funds the next model, which attracts more developers. That’s the flywheel.
OpenAI replicated it with Codex. It works. Now SpaceX has purchased the same capability wholesale.
Grok Bot, launched in beta on August 11, is SpaceXAI’s answer to Claude Code and Codex. Each bot is independently hosted on cloud servers and connects with your current tools to automate processes. It only needs your approval for the ultimate step. The product combines Grok’s reasoning with Cursor’s development environment. SpaceXAI is not just targeting developers with this. It’s intended for individuals engaged in work that results in documents, analyses, or presentations, as these can all be created programmatically, even if the person responsible for them doesn’t write any code themselves.
And alongside all of this, Imagine 2.0 shipped the same week. Not a video model; that stays at Imagine Video 1.5. But the image generation and editing capabilities are genuinely remarkable. Targeted editing with up to five reference images combined in a single generation. At launch, it ranked first globally in text-to-image and second in image editing on LMArena’s blind-vote leaderboard, just behind GPT Image 2. I’ve used it. The quality jump from the previous version is the kind where you stop and actually pay attention.
The Irony That Should Keep Anthropic Up at Night
And now the detail that makes this entire story structurally uncomfortable for one of the other frontier labs.
Anthropic currently pays SpaceX $1.25 billion per month to use Colossus 1. The contract runs through May 2029, totaling approximately $45 billion. Every Claude query processed during that contract term routes through SpaceX infrastructure. When the deal was signed in May, Grok was irrelevant. SpaceX had excess compute capacity and a competitor willing to pay handsomely for it.
Today, SpaceXAI owns Cursor, ships models that compete directly with Claude, sells a coding agent into the same market Anthropic targets, and is the infrastructure provider that Anthropic depends on training and run its models. The competitor is literally paying for the server time that helps build the tools competing with it. Neither party can easily walk away. The contract runs for three more years. But the strategic asymmetry has changed dramatically since May, and if I were running operations at Anthropic, this would be the slide that kept me awake.
Musk has already announced that Grok 4.7 is in progress. Training is reportedly complete. As part of the supplemental training phase, Cursor data is being combined with internal engineering data from SpaceX. He said publicly it will be “significantly better” and arrive in three to four weeks. Whether that timeline holds is anyone’s guess. But the direction isn’t.
Three Flywheels, One Race
When you step back far enough from the week’s noise, a structural pattern becomes obvious. The winning formula in AI in 2026 is code. Not code in the narrow technical sense, but code as the flywheel that generates data, revenue, and competitive advantage.
Anthropic figured this out first with Claude Code. OpenAI replicated it with Codex. SpaceX just bought it for $60 billion with Cursor.
Three frontier labs.
Three data flywheels.
Three loops where developers generate training data by using the product, which improves the model, which generates revenue, which funds the next iteration.
The consequence for the rest of us is straightforward: the models are getting better, faster, and cheaper. Grok 4.6 is 60 percent cheaper per token than its equivalents. Grok 4.7 is weeks away. Claude and OpenAI will respond with their own iterations. The competition compresses pricing and expands capability simultaneously.
If your work involves producing documents, analysis, code, presentations, research, or communication of any kind, the tools I’ve just described are not a future trend. They exist, they function, they cost less than they did last month, and in six months they will cost less again and do more. The question stopped being whether these tools would reshape how knowledge work happens. That part is settled. The question is whether you’ll have figured them out before knowing that knowing them stops being an advantage and starts being a minimum requirement.
I appreciate your time reading this. Follow me, comment your opinion, and don’t forget to subscribe to or support my newsletter for early access content.




