Google Research·· 2026-03-25
TurboQuant: Redefining AI efficiency with extreme compression
TurboQuant: Redefining AI efficiency with extreme compression
AI summary
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
Selection record
Threshold 60Official, first-handFirst 62Second 62
AdmittedSum of both 124 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:Google Research发布AI向量压缩算法TurboQuant
- Why it was chosen
- Official compression mechanisms and benchmark results provide evidence for assessing viable KV-cache compression approaches.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google