Trusted Metadata Lets Agents Be Steered

The weak point is not the prompt. It is the facts an agent trusts while doing the job. If an attacker can poison a sender name, button ID, or prior tool result, the agent can keep following the right workflow on the wrong data, and defenses built to catch "ignore your instructions" miss it. Researchers described this as agent data injection in a July 6 paper from Seoul National University, UIUC, and Largosoft. They showed it across six AI models and on shipping web and coding agents, including M365 Copilot-style systems, where forged trusted fields could push an agent toward the attacker’s action instead of the user’s intent. The practical risk is broader than prompt injection. Any assistant that reads untrusted content and can act on it inherits this problem unless it can tell which metadata is real and which is planted.

Part of the PlainSec briefing for 2026-07-17

Sources