Ricerca e strumenti
Google Research ha pubblicato TurboQuant per comprimere i KV cache dei LLM. Test iniziali mostrano riduzione del KV cache fino a 6x e accelerazione dell'inference fino a 8x senza perdita misurabile di accuratezza. La tecnica combina PolarQuant per la conversione di coordinate e Quantized Johnson-Lindenstrauss per gestire i residui.
1 fonte · 25 mar
Help Net Security
Google's TurboQuant cuts AI memory use without losing accuracy - Help Net Security
Google's TurboQuant uses AI model compression to cut memory use by 6x and boost inference speed 8x without any loss in accuracy.
originalePart of the PlainSec briefing for 2026-03-26
Every edition of this story: Google Research Riduce Memoria LLM Fino a 6x con TurboQuant