CrowdStrike says the most advanced publicly deployed content safety classifier, used to guard models such as Claude Opus 5.5 and Fable 5, can be bypassed by breaking a harmful request into harmless pieces. In testing, the team says the method worked across 9 of 10 offensive security categories.
The key weakness is sequence-level intent. The classifier checks each prompt on its own, so it can approve every subtask while attackers collect safe outputs and combine them outside the classifier’s view into something harmful. CrowdStrike says its direct attack tests against the guardrail otherwise produced a 0% bypass rate.
For teams that rely on per-request filtering as the main abuse control, the exposure is in the workflow, not the single prompt. If your safety layer only judges isolated messages, benign-looking decomposition can still defeat the policy even when no individual request looks malicious.
How Attackers Can Bypass LLM Safety Classifiers | CrowdStrike
The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier and found it can be systematically circumvented.