Google Research published TurboQuant to compress large language model key-value caches. Early tests show up to 6x reduction in KV cache memory and up to 8x faster inference with no measurable accuracy loss. The method pairs PolarQuant coordinate conversion with a Quantized Johnson‑Lindenstrauss residual step.
Part of the PlainSec briefing for 2026-03-26