Back to Papers & Articles

TurboQuant: Redefining AI efficiency with extreme compression

Amir Zandieh, Vahab Mirrokni — Google Research

Summary

TurboQuant is a compression algorithm that drastically reduces the size of high-dimensional vectors used in AI models. It combines PolarQuant and Quantized Johnson-Lindenstrauss methods to achieve high reduction in model size with near-zero accuracy loss — enabling faster vector search and reducing memory bottlenecks in LLMs without sacrificing performance.

Why it caught my eye

(notes to add)