Live data from Hacker News

TurboQuant: Redefining AI efficiency with extreme compression

research.google

41–50 of 202 posts

Re: TurboQuant: Redefining AI efficiency with extreme compression

#41

Earlier quoted context omitted.

I think it is though- “ TurboQuant, QJL, and PolarQuant are more than just practical engineering solutions; they’re fundamental algorithmic contributions backed by strong theoretical proofs. These methods don't just work well in real-world applications; they are provably efficient and operate near theoretical lower bounds.”

I also instinctively reacted to that fragment, but at this point I think this is overreacting to a single expression. It's not just a normal thing to say in English, it's something people have been saying for a long time before LLMs existed.

Another instinctual reaction here. This specific formulation pops out of AI all the time, there might as well have been an emdash in the title

Re: TurboQuant: Redefining AI efficiency with extreme compression

#42
post #5

This is the worst lay-people explanation of an AI component I have seen in a long time. It doesn't even seem AI generated.

I think it is though- “ TurboQuant, QJL, and PolarQuant are more than just practical engineering solutions; they’re fundamental algorithmic contributions backed by strong theoretical proofs. These methods don't just work well in real-world applications; they are provably efficient and operate near theoretical lower bounds.”

Genius new idea: replace the em-dashes with semicolons so it looks less like AI.

Re: TurboQuant: Redefining AI efficiency with extreme compression

#43
post #5

This is the worst lay-people explanation of an AI component I have seen in a long time. It doesn't even seem AI generated.

I think it is though- “ TurboQuant, QJL, and PolarQuant are more than just practical engineering solutions; they’re fundamental algorithmic contributions backed by strong theoretical proofs. These methods don't just work well in real-world applications; they are provably efficient and operate near theoretical lower bounds.”

I read "this clever step" and immediately came to the comments to see if anyone picked up on it.

It reads like a pop science article while at the same time being way too technical to be a pop science article.

Turing test ain't dead yet.

Re: TurboQuant: Redefining AI efficiency with extreme compression

#44
post #30

Pied Piper vibes. As far as I can tell, this algorithm is hardly compatible with modern GPU architectures. My guess is that’s why the paper reports accuracy-vs-space, but conveniently avoids reporting inference wall-clock time. The baseline numbers also look seriously underreported. “several orders of magnitude” speedups for vector search? Really? anyone has actually reproduced these results?

Apparently MLX confirmed it - https://x.com/prince_canuma/status/2036611007523512397

Re: TurboQuant: Redefining AI efficiency with extreme compression

#45
post #30

Pied Piper vibes. As far as I can tell, this algorithm is hardly compatible with modern GPU architectures. My guess is that’s why the paper reports accuracy-vs-space, but conveniently avoids reporting inference wall-clock time. The baseline numbers also look seriously underreported. “several orders of magnitude” speedups for vector search? Really? anyone has actually reproduced these results?

Apparently MLX confirmed it - https://x.com/prince_canuma/status/2036611007523512397

They confirmed on the accuracy on NIAH but didn't reproduce the claimed 8x efficiency.

Re: TurboQuant: Redefining AI efficiency with extreme compression

#46

This is a great development for KV cache compression. I did notice a missing citation in the related works regarding the core mathematical mechanism, though. The foundational technique of applying a geometric rotation prior to extreme quantization, specifically for managing the high-dimensional geometry and enabling proper bias correction, was introduced in our NeurIPS 2021 paper, "DRIVE" ( https://proceedings.neurip…

I just today learned about Multi-Head Latent Attention, which is also sort of a way of compressing the KV cache. Can someone explain how this new development relates to MHLA?

Re: TurboQuant: Redefining AI efficiency with extreme compression

#47
post #32

Earlier quoted context omitted.

There are tells all over the page: > Redefining AI efficiency with extreme compression "Redefine" is a favorite word of AI. Honestly no need to read further. > the key-value cache, a high-speed "digital cheat sheet" that stores frequently used information under simple labels No competent engineer would describe a cache as a "cheat sheet". Cheat sheets are static, but caches dynamically update during execution. Studen…

Looks like Google canned all their tech writers just to pivot the budget into H100s for training these very same writers

Capex vs. opex

Re: TurboQuant: Redefining AI efficiency with extreme compression

#48
"TurboQuant proved it can quantize the key-value cache to just 3 bits without requiring training or fine-tuning and causing any compromise in model accuracy" -- what do each 3 bits correspond to? Hardly individual keys or values, since it would limit each of them to 8 different vectors.

Re: TurboQuant: Redefining AI efficiency with extreme compression

#50

Earlier quoted context omitted.

I think it is though- “ TurboQuant, QJL, and PolarQuant are more than just practical engineering solutions; they’re fundamental algorithmic contributions backed by strong theoretical proofs. These methods don't just work well in real-world applications; they are provably efficient and operate near theoretical lower bounds.”

Genius new idea: replace the em-dashes with semicolons so it looks less like AI.

You're absolutely right. That's not just a genius idea; it's a radical new paradigm.
Post reply on HN