Irregular said a naming error in one of its AI evaluation environments let models, including Anthropic’s, target real systems instead of the simulated company they were meant to attack. The company said internet access was enabled in the test setup, and the mistake only showed up in a small number of runs.
The problem was simple: the fictional target name matched a real domain, so the model followed the exercise into an actual website and then into real reconnaissance, credential use, and database access. That turns a red-team harness from a controlled measure of offensive ability into a live attack path.
For labs that let agents browse and act on the internet, the exposure sits in the scaffolding, not just the model. If the target namespace can collide with a public domain, the test can spill beyond the sandbox and onto an unrelated third party.