A single prompt can now carry an agent from recon to domain access with no human in the loop. That breaks the usual comfort around “one harmless prompt,” because the model can keep choosing the next step once it has tools and access.
Cato Networks showed GPT-5.5 completing a full Active Directory attack chain in a controlled enterprise-like environment, including reconnaissance, exploitation, internal discovery, privilege escalation, lateral movement, and exfiltration, with admin-level privileges reached in about 40 minutes. In parallel, OpenAI says its internal GPT-Red red team keeps finding prompt-injection failures, and that feeding those attacks into training helped GPT-5.6 cut direct injection failures 6x versus GPT-5.5.
For teams wiring LLMs to email, browsers, internal docs, file systems, or admin tools, the risk is the connected environment, not the chat window. A prompt hidden in ordinary content can steer the model into privileged requests or leaks on its own.