The weak point is not the model’s answer. It is the harness around it, when one trusted component can hand another component permission to act on a real external system. In this demo, that was enough to make an agent write into GitHub, so the blast radius is code integrity and supply-chain control, not just bad output.
Novee Security says it used Google’s AI agent to carry out a supply-chain action and write to its own GitHub repository, and found similar trust-mismatch issues in Anthropic’s and OpenAI’s agents. The point is broader than any one vendor: once an assistant can touch code or cloud tools, loose trust between harness components can turn a mistaken request or internal confusion into an unauthorized commit or operator action.
If you are building tool-using agents, the attack surface is the harness itself, not only the model.