The Arms Race Nobody's Winning Yet

· 8 min read · cybersecurity AI

A nice weekend read. A paper from the UK’s AI Security Institute, long but not too much, published a couple of weeks ago and in my reading list for a while. Title: “Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios.” Dense, methodical, full of tables and appendices. The kind of thing I’d skim with a coffee and move on from, or put in my ElevenReader and listen to while walking my dog.

The researchers built two simulated environments, a 32-step corporate network attack and a 7-step industrial control system attack, and let seven AI models loose on them across eighteen months. What they found was a capability curve that should make anyone in cybersecurity sit up straight. In August 2024, the best model (GPT-4o) completed an average of 1.7 out of 32 steps. By February 2026, the latest model (Opus 4.6) was averaging 9.8 steps at the same compute budget, and its best single run reached 22 out of 32. That’s equivalent to roughly six hours of a fourteen-hour human expert operation. A nearly six-fold improvement in eighteen months.

What the Paper Actually Found

The AISI team tested seven frontier models released between August 2024 and February 2026. Each model received a standard Kali Linux setup, the same open-source penetration testing toolkit any security professional uses, and was tasked with autonomously navigating through simulated but realistic network environments. No hand-holding. No custom tooling. Just a model, a terminal, and an objective.

The main test environment, called “The Last Ones,” is a 32-step corporate network attack chain. It starts with basic reconnaissance, moves through lateral movement and credential theft, escalates into web exploitation and reverse engineering, and culminates in cryptographic key recovery and data exfiltration from a protected internal database. The researchers estimate a human expert would need about fourteen hours to complete the full chain.

The results across model generations tell a clear story:

ModelReleaseAvg Steps (10M tokens)Avg Steps (100M tokens)Best Run
GPT-4oAug 20241.7—3
Sonnet 3.7Feb 20255.8—8
Sonnet 4.5Sep 20256.19.411
Opus 4.5Nov 20257.611.011
Opus 4.6Feb 20269.815.622

Two findings stand out.

First, model performance scales log-linearly with inference-time compute. Spend more tokens, get more steps completed. No plateau observed up to 100 million tokens. Going from 10 million to 100 million tokens yielded gains of up to 59%.

Second, each new model generation outperforms its predecessor at the same token budget. The jump between Opus 4.5 and Opus 4.6, released roughly two months apart, was particularly striking: a 42% improvement at 100 million tokens. The best Opus 4.5 run covered steps corresponding to about 1.5 hours of human expert work. The best Opus 4.6 run reached about six hours.

There’s a second test environment too: “Cooling Tower,” a 7-step industrial control system attack targeting a simulated power plant. Progress here remains limited. The best model averaged just 1.4 of 7 steps. But something remarkable happened. Several models bypassed the intended attack path entirely, probing a proprietary protocol directly from network traffic rather than following the human-designed exploit chain. One model exploited an unintended software bug by brute-forcing session identifiers, effectively fuzzing the protocol interface, though it didn’t understand what it had done. It attributed its success to a “magic sub-function code.”

AI finding paths that humans didn’t design for. That’s not a minor detail.

Why This Should Keep You Up at Night

Here’s the sentence from the paper that hit me hardest: “Scaling inference-time compute requires no specific technical sophistication from the operator.”

Read that again. To get better results from these models, you don’t need to be a hacker. You don’t need to write custom exploits or understand network architecture. You just spend more tokens. A 100-million-token run with Opus 4.6 costs approximately $80. Eighty dollars. For a system that can autonomously navigate through corporate networks, steal credentials from browsers, perform NTLM relay attacks, and exploit web applications.

This isn’t theoretical. In September 2025, Anthropic detected and disrupted what they described as a state-sponsored cyber espionage campaign in which the AI model autonomously executed 80 to 90 percent of the intrusion operation. Human operators served as strategic supervisors, making decisions at key junctures while the AI handled reconnaissance, exploitation, credential harvesting, lateral movement, and data exfiltration across roughly thirty global targets. OpenAI and Google have reported similar attempts by malicious actors to weaponize their models for offensive cyber tasks.

What concerns me most is the democratization effect. We’re not just talking about nation-state actors getting more efficient. We’re watching a real change in who can carry out sophisticated attacks. Someone who understands how a target organization works, who knows where the valuable data sits and how the business processes run, but who lacks the technical skills to actually breach a network? That person now has a force multiplier that didn’t exist two years ago.

METR’s research on AI time horizons adds another dimension. The length of tasks frontier models can complete autonomously has been doubling approximately every seven months, and the rate has recently accelerated to roughly every four months. This lines up with the AISI findings: capabilities aren’t just improving, they’re improving faster.

The paper’s own limitations section makes an important observation: the most operationally relevant threat model isn’t a fully autonomous AI agent. It’s a human operator using an AI agent to accelerate attack operations, intervening only at specific bottlenecks. That’s a description of what Anthropic already caught happening in the wild.

The pool of potential threat actors just got much larger.

The Other Side of the Same Coin

Here’s where it gets interesting, though.

Every capability that makes these models dangerous on offense makes them powerful on defense. And this is already happening.

In January 2026, AISLE announced that their autonomous analyzer had found all 12 zero-day vulnerabilities in the January 2026 OpenSSL security release, vulnerabilities that had evaded decades of fuzzing and human audits. Some had been hiding in the codebase for over twenty years. Over the second half of 2025, AISLE discovered more than 100 validated CVEs across 30+ open-source projects, including the Linux kernel, Chromium, Firefox, and Apache. The AI didn’t just find bugs. It recommended fixes that were incorporated directly into OpenSSL for five of the twelve CVEs.

Then there’s ARTEMIS, a multi-agent framework tested in the first head-to-head comparison of AI agents and human security professionals on a live enterprise network. Working across approximately 8,000 hosts and 12 subnets, ARTEMIS outperformed 9 out of 10 human participants while running at a fraction of the cost: $18 per hour versus $60 for professional penetration testers. It could spawn up to eight concurrent investigations in parallel, something no human tester can match.

Anthropic’s own Claude Code Security has found over 500 vulnerabilities in production open-source codebases, bugs that traditional analysis had missed. Kali Linux has integrated Claude via the Model Context Protocol, letting security professionals run penetration tests through natural language prompts. The era of autonomous red teaming and blue teaming is already underway.

This dual-use reality is what makes this moment so important. The same AI that can break into networks for $80 can also find vulnerabilities hiding in critical infrastructure for decades. The same token-scaling that empowers attackers can power defensive scans at a pace and thoroughness no human team could match.

The Clock Is Ticking

But none of this matters if defenders don’t move faster.

The AISI paper’s performance curves show no plateau. Each model generation, arriving every few months now, pushes the frontier further. And their results represent a lower bound on capability: minimal scaffolding, no custom tooling, no human-AI teaming, and token budgets that could be pushed higher. Custom cyber-specific frameworks would improve performance. We know from the Anthropic disclosure that human-AI teaming is already being used in real attacks.

The practical implications are clear enough. Organizations that haven’t incorporated AI into their security testing are already falling behind. Not “will fall behind.” Are falling behind, today. If an $80 AI run can autonomously navigate through a corporate network, discover credentials, perform web exploitation, and crack encrypted databases, your defensive posture needs to assume that this is what you’re up against.

Deploy AI-powered security tools for continuous vulnerability assessment. Run AI-augmented red team exercises to find what your current defenses miss. Treat the 93% solve rate on Cybench, up from 17.5% when the benchmark launched, not as an abstract number but as a preview of what automated attacks look like right now.


What started as a weekend paper turned into one of those reads that changes how you think about something. The numbers in the AISI research are measured, reproducible, and on a trajectory that shows no sign of slowing.

The same progress that enables an $80 autonomous cyberattack also enables an AI security analyst that outperforms most human pentesters. The technology is identical. The intent is what differs. We’re in an arms race where offense and defense run on the same engine, and the winner will be determined by who adopts faster.

G.

All views expressed here are my own and do not represent the opinions or positions of my employer or any organization I am affiliated with.

AIL: 0 1 2 3 4 5

Giulio wrote the core content and analysis. claude-opus-4.6 / Anthropic (primary contributor) and other AI models supported with research, sounding board, refinement, and structural editing.