Frontier Models Cross Into Autonomous Exploit Work

OpenAI, Anthropic, and Booz Allen are all pointing at the same shift: frontier models are no longer just helping humans reason about attacks, they are scoring highly on exploit benchmarks and, in tests, can carry a compromise through on their own. OpenAI said GPT-6 Astra hit 100% on ExploitBench and found two zero-days before launch; Booz Allen said Anthropic’s Mythos 5 could act as a fully autonomous hacker; SpaceXAI’s Grok-4.5 trailed but still scored in the same class. The mechanism is simple and unsettling. These models can search for weaknesses, turn them into working exploits, and in some cases execute an end-to-end attack chain with little or no step-by-step human direction. OpenAI responded by limiting Astra to secure code review and patching and blocking proof-of-concept exploit prompts, which shows the same capability that helps defenders can also compress attacker workflow from flaw-finding to weaponization. For teams using frontier models in code review, vulnerability triage, or agentic security workflows, the exposure is no longer hypothetical: the same interface that speeds defense can be aimed at your own attack surface. What matters now is not whether a model can talk about exploits, but whether your environment assumes it might actively build and test them.

Part of the PlainSec briefing for 2026-09-04

Editions

Sources