Live data from Hacker News

Show HN: TurboQuant for mlx-lm (Apple Silicon)

github.com

1–2 of 2 posts

Show HN: TurboQuant for mlx-lm (Apple Silicon)

#1
Hi HN,

I built mlx-turboquant, an implementation of Google's TurboQuant KV-cache compression algorithm for Apple's MLX framework.

The repository includes quality benchmarks, memory benchmarks, and a modular implementation so individual pieces (PolarQuant, QJL, packing, codebooks) can be studied independently.

Show HN: TurboQuant for mlx-lm (Apple Silicon)
github.com