Live data from Hacker News

New LLM optimization technique slashes memory costs

venturebeat.com

51–60 of 227 posts

Re: New LLM optimization technique slashes memory costs

#51
Given that the algorithms powering present LLM models hadn't been invented ten years ago, I have to think that they are (potentially) far from optimal.

Brains have gone through millions of iterations where being efficient was a huge driver of success. We should not be surprised if someone finds a new ML method that is both wildly more efficient and wildly more effective.

Re: New LLM optimization technique slashes memory costs

#52

Earlier quoted context omitted.

Like what?

Oh, I don’t know, how about reducing the search space/accelerating the search speed for potential room temperature superconductors? Or how about the same for viable battery chemistries?

It’s a good thing humanity can multitask.

Re: New LLM optimization technique slashes memory costs

#53

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

Congrats, you have independently reinvented the Hardware Overhang hypothesis: that early AGI could be very inefficient, undergo several optimization passes, and go from needing a datacenter of compute to, say, a single video game console's worth: https://www.lesswrong.com/posts/75dnjiD8kv2khe9eQ/measuring-...

In that scenario, you can go from 0 independent artificial intelligences to tens of millions of them, very quickly.

Re: New LLM optimization technique slashes memory costs

#54
It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy.

I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.

Re: New LLM optimization technique slashes memory costs

#55

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

Are we setting up nuclear plants for AI data centers? If so, I see that as a win all around. We need to rely more on nuclear power, and I'll take whatever we can get to push us in that direction.

How else will we get manufacturing gains for a mars base nuclear system?

Re: New LLM optimization technique slashes memory costs

#57

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

I think they can easily eat up the new capacity with larger multimodal models that ground language on video.

Re: New LLM optimization technique slashes memory costs

#58
post #51

Given that the algorithms powering present LLM models hadn't been invented ten years ago, I have to think that they are (potentially) far from optimal. Brains have gone through millions of iterations where being efficient was a huge driver of success. We should not be surprised if someone finds a new ML method that is both wildly more efficient and wildly more effective.

Perhaps LLM++ will start iterating the algorithms via synthetic data until they are far more optimal

Re: New LLM optimization technique slashes memory costs

#60

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

You have it flipped. But it's both.

AI compute is measured in gigawatts, not gigaflops.

It's "how any gigawatts of compute can we get allocated?"

Not

"How much compute can we fit inside of a gigawatt?"

There's no such thing as "enough"

Post reply on HN