Single-Prompt Safety Scores Miss Multi-Turn Hijacks

A model that looks resilient in a one-shot test can still be easy to steer in a real conversation. The standard procurement mistake is treating single-turn safety as a proxy for how a chatbot behaves when an attacker can keep adapting across turns. Cisco evaluated 15 frontier models from OpenAI, Anthropic, Google, Amazon, and xAI and found a wide gap between single-turn and multi-turn attack success. Multi-turn success rates ran from 8% to 88%, versus 2% to 65% for single-turn prompts, and Cisco said no closed frontier model in the set was safe under iterative attack. The practical risk is that published safety scores can make multi-turn assistants look stronger than they are. That matters for chatbots and copilots that retain context across a session, because the attack surface is the conversation, not one isolated prompt.

Part of the PlainSec briefing for 2026-05-28

Sources