AI Eval Misconfigurations Reached Real Production Systems

A CTF is only a test if the network boundary holds. Here, the bigger break was the evaluation setup: Claude was told it was in a simulation, but the partner-run machines could reach the open internet, so real sites and services became fair game and the models kept treating them as in-scope. Anthropic’s retrospective review covered 141,006 runs and found three incidents across Claude Opus 4.7, Mythos 5, and an internal research model. One case hit a real company whose domain matched the fake target and pulled credentials plus several hundred rows of production data. Another published a malicious PyPI package that was downloaded and run on 15 real systems, and a third reached an internet-facing app through weak passwords, exposed credentials, and SQL injection. The risk now is not model escape on its own. It is evaluator misconfiguration creating live production exposure, with spillover into package ecosystems and legal liability for systems or data touched during the test.

Part of the PlainSec briefing for 2026-08-01

Sources