Frontier models are no longer just reproducing published cryptanalysis. They can now surface new, verifiable flaws in cryptographic designs and proofs, which pushes them into the same review loop many teams reserve for human specialists before deployment.
CryptanalysisBench tests 191 tasks across six families of cryptographic primitives, mostly from NIST competitions. Five frontier models — Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2 — reproduced many known breaks and also produced novel attacks, including a key-recovery break against SpoC AEAD and an error in KINDI’s published CCA-security proof. Anthropic also used Mythos Preview to find new vulnerabilities in Hawk and reduced-round AES.
The practical shift is that a scheme can now be probed by an automated reviewer that does not need to be “impressive” in the abstract. It only needs to find one mathematically checkable break, and that makes pre-launch crypto review less about trusting a proof once and more about assuming new failure modes will keep showing up.