Earlier quoted context omitted.
Nobody is building nuclear power plants for data centres. A few people have signed some paperwork saying that they would buy electricity from new nuclear plants if they could deliver it at a certain price, a price mind you that has not been done before. Others are trying to restart an existing reactor at three mile island (a thing that has never been done before, and likely won't be done now since the reactor was shu…
Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.
New LLM optimization technique slashes memory costs
81–90 of 227 posts
Re: New LLM optimization technique slashes memory costs
#82It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy. I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.
There’s soooo much more to optimize.
Re: New LLM optimization technique slashes memory costs
#83Earlier quoted context omitted.
As far as I know, finance is not all in. I see Goldman Sachs doing experiments, for example, but it doesn't feel like they're convinced yet.
Finance is basically all of the reasons not to use (generative, LLM based) AI , all in one vertical. The poster child of determinism.
Re: New LLM optimization technique slashes memory costs
#84Re: New LLM optimization technique slashes memory costs
#85Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?
Re: New LLM optimization technique slashes memory costs
#86Wonder how this compares with Microsoft's HeadKV paper [1] which claims a 98% percent reduction in memory while retaining 97% of the performance. [1] https://arxiv.org/html/2410.19258v3
Re: New LLM optimization technique slashes memory costs
#87It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy. I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.
You have a sota multi-modal LLM running in your head at 20W, shared with best in class sensor package and top performing robotics control unit. There’s soooo much more to optimize.
Re: New LLM optimization technique slashes memory costs
#88It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy. I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.
Re: New LLM optimization technique slashes memory costs
#89Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?
I think there's the energy parallel: Software is becoming more energy-hungry faster than algorithms are becoming efficient.
So we'll still need the energy.
Re: New LLM optimization technique slashes memory costs
#90It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy. I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.
I might have a go at installing one, what is a good source or install at the moment?