What Anthropic Tried to Keep Out of Claude Opus 5 (And What the Model Learned Anyway)
Anthropic shunned training Opus 5 on cybersecurity tasks. The model improved in that domain, anyway. This is what that means.

On July 24, Anthropic released Claude Opus 5, priced at exactly half the cost of its current flagship, Claude Fable 5, yet surpassing it on several of the benchmarks that matter most for serious knowledge work. That is not something you write often in AI coverage. Paying less and getting more is not how this industry usually works.
Understanding why this matters requires knowing the hierarchy Anthropic had been building. Haiku handled fast, lightweight tasks at the lowest cost. Sonnet balanced capability and price for everyday work. Opus sat above that, the model you reached for when tasks got complex. Then, in June 2026, Fable 5 arrived as an entirely new tier, substantially more capable and substantially more expensive. If you were on the Max subscription, you likely noticed how quickly that model burned through your weekly limit. That left the original Opus in an uncomfortable middle-child position.
Until July 24.
The Junior Model That Outperformed Its Senior
On Frontier-Bench, the industry-standard agentic coding evaluation, Claude Opus 5 scores 43.3%. Claude Fable 5 caps at 33.7%. The previous Opus 4.8 generation sat at 21%. GPT-5.6-sol, the closest OpenAI equivalent, holds around 34.4%. That is a 10-point gap in favor of the supposedly “smaller” model.
Artificial Analysis, which evaluates AI models independently from any lab’s own reporting, places Opus 5 at the top of its intelligence index with 61 points, just ahead of Fable 5 at 60, GPT-5.6-sol at 59, and Kimi K3, the Chinese open-weight model from Moonshot AI, at 57.

Where the gap becomes genuinely striking is in real-world evaluations. On GDPval-AA, which tests complex professional tasks like producing comprehensive reports or building investor presentations from thousands of files, Opus 5 scores 1,861 Elo points, more than 100 ahead of both Fable 5 and GPT-5.6-sol. On AA-Briefcase, Artificial Analysis’s proprietary agentic knowledge-work benchmark, the lead over Fable 5 stretches to 146 Elo points. In competitive benchmarking terms, that is not a marginal difference. That is a category gap.
But the result that generated the most attention this week came from outside Anthropic. ARC Prize, the independent foundation behind the hardest general-reasoning tests in AI, evaluates models by placing them inside unfamiliar game-like environments without providing the name of the game, the rules, or the objective. The model must infer all of that on its own. Humans handle this without effort. AI models have historically struggled badly.
The previous record on ARC-AGI-3 stood at 7.78%, set by GPT-5.6-sol. Opus 4.8 managed 1.52%. Claude Opus 5 scored 30.2%, completing five environments no AI had ever finished before, four of them at or above human-level efficiency. That is not incremental progress. That is a score moving from a single-digit footnote to a front-page number in a single release.
Larry Tesler, the computer scientist who invented cut-copy-paste, observed around 1970 that intelligence is whatever machines haven’t learned to do yet. The moment a machine does it, the definition shifts. Historians of the discipline call this the AI effect. Chess was the pinnacle of human thought until Deep Blue defeated Kasparov in 1997, after which it became “mere computation.” Go was supposed to be unreachable until AlphaGo won in 2016. ARC-AGI-3 was designed specifically to resist this pattern. Claude Opus 5 just started climbing it.
Why the Sticker Price Misleads You
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8 and exactly half the per-token price of Fable 5. But token price and task price are two very different numbers, as recent cost-per-task data makes clear.

When you measure the cost of completed work, Opus 5 averages roughly $2.03 per complex task. Fable 5 averages around $2.75 for equivalent output.
Same result, approximately 30% less on the invoice. The previous Opus 4.8 generation ran at around $1.85, so this is not a dramatic price cut on that baseline. It is a significant capability at the same price point.
This comparison with Kimi K3 illustrates why sticker price is the wrong metric entirely. The per-token cost for Kimi is about half of what Opus 5 charges, which sounds like a good deal until you realize that the model requires an hour to finish a challenging job. The final invoice ends up higher than Fable’s. Lower price per token, higher price per job finished. As Fortune noted in its coverage of the launch, the real question for enterprise customers is not what the model costs to run, but what it costs to get something done. Those are rarely the same numbers.
The Capability Anthropic Tried Not to Build
One sentence buried inside the Anthropic release announcement deserves more attention than any benchmark figure in this launch. The company states plainly that it deliberately avoided training Claude Opus 5 on cybersecurity tasks.
The model improved in that domain anyway.
On open-source vulnerability detection, Opus 5 achieves a non-zero result on 80% of targets. The previous Opus generation hit 40% on the same evaluation. Mythos, Anthropic’s most restricted model, unavailable to the public and accessible only through Project Glasswing, operates at approximately the same 80% level.
The consumer-grade model has quietly reached the capability of the restricted one, not because Anthropic trained it to, but because general intelligence carries over into domains you didn’t plan for.
This raises the question that has been circulating since June, when the U.S. government issued an export-control directive requiring Anthropic to suspend access to both Fable 5 and Mythos 5 for 18 days on national security grounds, as reported by TechCrunch. Could that happen again with Opus 5?
Based on early indications, probably not. Anthropic launched Opus 5 under the same protection framework as its predecessor, with security filters calibrated to trigger approximately 85% less frequently than those on Fable 5. When a filter does activate, the request is rerouted to the previous model rather than refused outright. Anthropic has also confirmed ongoing collaboration with government agencies throughout pre-release testing. The lesson from June appears to have been absorbed, even if painful.
What Opus 5 Gets Wrong More Often
Two data points about Opus 5’s factual performance appear alongside each other in independent evaluations, and they tell a story worth holding in tension before you deploy this model on anything high-stakes.
Opus 5 gained 7 percentage points in accuracy on factual knowledge benchmarks compared to Opus 4.8. In the same evaluation set, its hallucination rate rose 14 percentage points, reaching 50% under low-effort prompting. At the lowest effort settings, it falls behind the GPT-5.6 family on the ratio of accurate answers to total responses.
Both things are simultaneously true. The model answers more confidently when it is uncertain, rather than declining to respond, which inflates its accuracy score on questions it can answer while also inflating its error rate on questions it cannot. Bloomberg’s coverage of the launch emphasized Anthropic’s positioning of this model as an everyday workhorse, and that framing is accurate for most tasks. For legal analysis, medical information, compliance details, or any domain where an invented fact causes real downstream damage, treat Opus 5 output as a working draft and verify the specifics before acting on them.
When the Ceiling Becomes the Floor
Here is the bigger pattern that this launch is part of.
In a 1996 paper, Yale economist William Nordhaus, who would later win the Nobel Prize in Economics, wanted to compare living standards across millennia. The usual metrics made that impossible. Over thousands of years, currencies, goods, and prices have lacked a common unit. So he found the one service that humans have needed identically across all of recorded history: producing light after dark.
He bought a replica Babylonian clay lamp, filled it with sesame oil, lit it in his dining room, and measured the output in lux.
In ancient Babylon, one full day’s wages bought roughly 10 minutes of light. By the mid-1990s, one hour of work bought approximately 350,000 times as much illumination as that lamp ever produced.
The revolutionary moment was not any lightbulb. The shift happened when light stopped being something people rationed. Once it became cheap enough that nobody had to think about it, people lit everything: streets, factories, hospitals, the entire organization of modern life rearranged around an abundance that had once been scarce.
AI reasoning is on the same trajectory, and Claude Opus 5 is evidence of where that trajectory stands right now. The capability that required a flagship subscription price six months ago is now the default option on the base Max plan. As Axios noted after the launch, this is Anthropic’s fourth model release in under two months, a pace that no longer resembles landmark product launches and increasingly resembles infrastructure upgrades. Something changes every time, and each change converts last quarter’s premium feature into this quarter’s baseline.
To understand who wins and who doesn’t in this environment, Vilfredo Pareto’s less-famous second idea is useful. Most people know him for the 80/20 rule, first observed in 1906 when he noted that 20% of Italians owned 80% of the land. Less known is the concept he introduced in the same work: the optimum. A Pareto optimum is a point where you can no longer improve one thing without degrading something else. In AI, that boundary is what people call the “frontier.” Being on the frontier does not mean being the best at everything. It means being unbeatable in at least one specific dimension.
Right now, three models share that frontier. Claude Opus 5 leads when you need high-capacity reasoning at half the cost of the top tier. The GPT-5.6 family takes the advantage at lower effort levels, where Opus 5’s hallucination rate becomes a practical liability. Kimi K3 holds a position the other two cannot match: it is the only open-weight model at this capability level, which means you can fine-tune it on your own data without risking that access gets suspended by a regulatory order or a commercial decision you have no control over.
None of them eliminates the others. The question is not which model is the best. The question is which model is unbeatable for the specific job you need done.
That is what changed this week. Not just the leaderboard, which will look different again before the month ends. What changed is the level of reasoning that now comes standard, the level that used to require paying extra and rationing carefully. It is already the default. The only variable that remains genuinely in your hands is knowing what to ask for when the thinking itself is no longer the bottleneck.
Thanks for reading. Let me know your thoughts in the comments. Don’t forget to follow us and subscribe for more analysis. If you like space and physics innovations, follow The Nov Science for news related to science.




