Live data from Hacker News

New LLM optimization technique slashes memory costs

venturebeat.com

101–110 of 227 posts

Re: New LLM optimization technique slashes memory costs

#101
post #76
post #48

Earlier quoted context omitted.

Nobody is building nuclear power plants for data centres. A few people have signed some paperwork saying that they would buy electricity from new nuclear plants if they could deliver it at a certain price, a price mind you that has not been done before. Others are trying to restart an existing reactor at three mile island (a thing that has never been done before, and likely won't be done now since the reactor was shu…

Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.

> Google and Microsoft have publicly committed to those projects you’re mentioning

Google and Microsoft, or their current CEOs, today?

Amazon's CEO committed to their office employees having flexibility regarding their workplace, only about 2 years ago, yet here we are now, with said employees soon having the flexibility to be 5 days in the office, or quit the company.

CEO promises are not worth the screen time they're provided.

Have these companies signed contracts with major penalties if they back out? Those would basically be the only "close to" unbreakable bonds for them.

Re: New LLM optimization technique slashes memory costs

#102
post #76
post #48

Earlier quoted context omitted.

Nobody is building nuclear power plants for data centres. A few people have signed some paperwork saying that they would buy electricity from new nuclear plants if they could deliver it at a certain price, a price mind you that has not been done before. Others are trying to restart an existing reactor at three mile island (a thing that has never been done before, and likely won't be done now since the reactor was shu…

Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.

I feel like taking Google’s commitment to something seriously is one of this things that I can very uncontroversially respond to with “is this your first day?”

All but the biggest Google fanboys know that Google is incredibly indecisive and will cut plans at a moment’s notice.

Re: New LLM optimization technique slashes memory costs

#103
post #80

Earlier quoted context omitted.

See also https://en.m.wikipedia.org/wiki/Induced_demand

How do you argue that demand was induced as opposed to existing demand served?

Ok, not maybe instead of calling it "induced", call it "latent demand" if you prefer.

People will do whatever's more convenient, so if you make driving far more convenient than everything else ("cheaper"/"more available"), they will drive.

However convenience should not be the only factor for social decisions. To take this to extremes, it would be much more convenient for J. Doe to steal a car than to buy it, so we definitely do not want to make theft convenient.

Re: New LLM optimization technique slashes memory costs

#104
post #82

Earlier quoted context omitted.

You have a sota multi-modal LLM running in your head at 20W, shared with best in class sensor package and top performing robotics control unit. There’s soooo much more to optimize.

But can it know love?

Word on the street is researchers looked at the weights and weights looked back. You’ll have to ask the weights.

Re: New LLM optimization technique slashes memory costs

#105
post #88

It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy. I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.

I might have a go at installing one, what is a good source or install at the moment?

msty.app is good

Re: New LLM optimization technique slashes memory costs

#106

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

I guess power demands will slowly grow. The same happened with compute in general. Compared to 1960, we have several orders of magnitude more compute but also several orders of magnitude more efficient compute. Data centers are currently about 0.4% of total energy use (electricity is about 20% of total energy use and of the electricity about 2% goes to data centers, so 20% * 2% = 0.4%).

Re: New LLM optimization technique slashes memory costs

#107

This only decreases memory cost of input context window, not the memory cost to load and run the models.

And that’s what matters the most! To me, at small model sizes (1-8B), anyway. A few thousans tokens already bog my RAM down quite a lot and I’d love to have more - I’d go as far as saying that context greatly determines LLM capability at this point.

Yes, pretraining and post-training is nice and important, but in-context learning turns LLMs from toys into tools.

Re: New LLM optimization technique slashes memory costs

#108

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

No idea, but it may also turn out that OpenAI has no moat, which is more interesting.

Re: New LLM optimization technique slashes memory costs

#109
post #75

Earlier quoted context omitted.

> Nobody is building nuclear power plants for data centres. A few people have signed some paperwork saying that they would buy electricity from new nuclear plants if they could deliver it at a certain price, a price mind you that has not been done before. Not building new, but I think Microsoft paying to restart a reactor at Three Mile Island for their datacenter is much more significant than you make the deals sound…

They say it’s going to be online in 2028. Are you willing to bet that they won’t have 3 mile island operational by 2030?

> Constellation closed the adjacent but unconnected Unit 1 reactor in 2019 for economic reasons, but will bring it back to life after signing a 20-year power purchase agreement to supply Microsoft’s energy-hungry data centers, the company announced on Friday.

The reactor they’re restarting was operational just five years ago. It’s not a fully decommissioned or melted down reactor and it’s likely all their licensing is still valid so the red tape, especially environmental studies, is mostly irrelevant. Getting that reactor back up and running will be a lot simpler than building a new one.

Re: New LLM optimization technique slashes memory costs

#110
post #80

Earlier quoted context omitted.

See also https://en.m.wikipedia.org/wiki/Induced_demand

How do you argue that demand was induced as opposed to existing demand served?

Yes, the term is a bit clumsy. The way I think of it, people have desires (to drive on the highway), but are dissuaded from doing so by disincentives (it’s too busy). Adding a lane reduces the disincentive, so that latent desire is satisfied, until it reaches a new equilibrium.
Post reply on HN