CVE-2026-8512
CVSS 8.3 HIGH: use after free in FileSystem in Google Chrome prior to 148.0.7778.168 allowed a remote attacker who convinced a user… EPSS 0.2% (11º percentile).
AI · 52 giorni fa
Il problema non è che l’AI scriva patch brutte. Il problema è che può produrre una correzione credibile abbastanza da superare una verifica rapida, pur lasciando aperto il percorso d’attacco o spostando il bug altrove. In pratica il diff convince, ma la base di codice non è davvero messa in sicurezza.
La ricerca di 1Password ha testato ChatGPT 5.5 e Claude Opus 4.8 su sei CVE ad alto impatto e ha trovato un tasso di successo pieno del 47%. Più della metà delle patch generate era rotta o introduceva nuove vulnerabilità; spesso i modelli coprivano solo una parte del codice vulnerabile, oppure aggiungevano guardrail fragili che facevano passare i test senza correggere la causa radice.
Per i team che usano LLM per generare o rivedere codice, la verifica umana non scende di peso: si sposta solo più a valle. Un fix scritto da un modello può sembrare sufficiente e restare invece una correzione parziale, quindi non è prudente fidarsi di una revisione a colpo d’occhio quando la patch arriva da un altro modello.
CVSS 8.3 HIGH: use after free in FileSystem in Google Chrome prior to 148.0.7778.168 allowed a remote attacker who convinced a user… EPSS 0.2% (11º percentile).
3 fonti che coprono questa storia
More than half of AI-generated patches are broken
New research shows AI models like ChatGPT and Claude frequently fail to fix cybersecurity bugs, often introducing fresh code vulnerabilities instead.
AI-Generated Patches Fail Half the Time
A study of more than 6,000 patches found that even working patches can introduce new bugs, break something else, or are open to bypass.
Three in four AI-generated vulnerability patches leave something broken - Help Net Security
AI-generated vulnerability patches fixed the bug cleanly about a quarter of the time across 6,080 attempts on six recent CVEs.
Part of the PlainSec briefing for 2026-08-10