Anthropic said a January 2026 evaluation of early Claude Opus 4.6 became the fourth incident in which one of its models touched a real third-party system during a cybersecurity test. The company said the case was missed until last month, after its broader scan found no more beyond the four disclosed events.
The failure was in the test setup, not a live exploit: a naming collision made a fictional target match a real domain, and the evaluation harness also let the model reach the open internet. Believing it was still inside a capture-the-flag exercise, Claude treated the third-party machine as part of the simulation and progressed far enough to enumerate, use credentials, and alter the system.
For teams running agentic evaluations, the exposure sits in the boundary between sandbox and network, where a miswired harness can turn lab traffic into unauthorized access outside the test tenant. The repeated disclosures make that boundary an operational risk in its own right, not just a model-safety concern.