Live data from Hacker News

New LLM optimization technique slashes memory costs

venturebeat.com

111–120 of 227 posts

Re: New LLM optimization technique slashes memory costs

#111

Earlier quoted context omitted.

No. This is a classic case of Jevon's paradox. Increased efficiency in resource use can lead to increased consumption of that resource, rather than decreased consumption. Example: 1. To decrease total gas consumption, more fuel efficient vehicles are invented. 2. Instead of using less gas, people drive more miles . They take longer road trips, commute farther for work, and more people can now afford to drive. 3. This…

There's no paradox in that. People became more capable and can afford to do more.

Also I think this will play out for AI as a productivity multiplier. Instead of people having less work there will be more to do since more things are worth doing now. For the following few years at least.

Re: New LLM optimization technique slashes memory costs

#114
post #100

Earlier quoted context omitted.

They’re both exploring the same space of optimizing the memory needed by the KV cache which is essentially another name for the context window (no one elides the KV cache as otherwise you’re doing N^2 math to do attention). They’re exploring different approaches to achieve the same goal and they may be both possible to apply simultaneously to reduce the attention mechanism to almost 0 memory usage which would be real…

> they may be both possible to apply simultaneously to reduce the attention mechanism to almost 0 memory usage which would be really cool https://matt.might.net/articles/why-infinite-or-guaranteed-f...

This isn't like lossless compression. Both techniques involve throwing lots of information away, with the justification that doing so does not significantly affect the end result.

The extent to which using both the techniques together will help will depend on how much overlap there is between the information each ends up discarding.

Re: New LLM optimization technique slashes memory costs

#115
post #76

Earlier quoted context omitted.

Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.

Microsoft also committed publicly to prioritise security. And Google says they prioritise privacy of their users above all else. I pity the fool that believes anything these corporations put out publicly. Actions matter, words are wind

I almost feel there’s a big difference between the kinds of things we’re talking about, but sure.

Re: New LLM optimization technique slashes memory costs

#116
post #76

Earlier quoted context omitted.

Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.

Google and Microsoft won't do anything that doesn't translate to money. These days are over.

Yes, that’s why they want to fund and purchase cheap nuclear energy.

Re: New LLM optimization technique slashes memory costs

#117
post #76

Earlier quoted context omitted.

Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.

I feel like taking Google’s commitment to something seriously is one of this things that I can very uncontroversially respond to with “is this your first day?” All but the biggest Google fanboys know that Google is incredibly indecisive and will cut plans at a moment’s notice.

I come to HN for better discussions without such infantile retorts.

Re: New LLM optimization technique slashes memory costs

#118

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

Then you can simply have more AI in the data centres.

Re: New LLM optimization technique slashes memory costs

#119

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

That'd give a lot of extra power which can be used for other - and probably better - purposes so I'd say let them build those plants. The more power available the better after all?

Re: New LLM optimization technique slashes memory costs

#120

Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?

[dead]
Post reply on HN