AI · 110 days ago
A model that looks resilient in a one-shot test can still be easy to steer in a real conversation. The standard procurement mistake is treating single-turn safety as a proxy for how a chatbot behaves when an attacker can keep adapting across turns.
Cisco evaluated 15 frontier models from OpenAI, Anthropic, Google, Amazon, and xAI and found a wide gap between single-turn and multi-turn attack success. Multi-turn success rates ran from 8% to 88%, versus 2% to 65% for single-turn prompts, and Cisco said no closed frontier model in the set was safe under iterative attack.
The practical risk is that published safety scores can make multi-turn assistants look stronger than they are. That matters for chatbots and copilots that retain context across a session, because the attack surface is the conversation, not one isolated prompt.
3 sources covering this story
Frontier AI models collapse under multi-turn AI attacks, Cisco finds - Help Net Security
Cisco research finds multi-turn AI attacks push success rates as high as 88% across 15 flagship models from OpenAI, Anthropic, Google, xAI.
AI models more vulnerable than claimed when faced with iterative attacks
Cisco researchers show how leading AI models wither under realistic multi-turn attacks, calling into question the value of vendors’ single-prompt safety benchmarks.
Leading AI models are more vulnerable to malicious prompts than vendors claim
Hackers could subvert frontier models with attacks that their developers overlook, Cisco said.
Part of the PlainSec briefing for 2026-05-28