Earlier quoted context omitted.
No. This is a classic case of Jevon's paradox. Increased efficiency in resource use can lead to increased consumption of that resource, rather than decreased consumption. Example: 1. To decrease total gas consumption, more fuel efficient vehicles are invented. 2. Instead of using less gas, people drive more miles . They take longer road trips, commute farther for work, and more people can now afford to drive. 3. This…
There's no paradox in that. People became more capable and can afford to do more.
New LLM optimization technique slashes memory costs
111–120 of 227 posts
Re: New LLM optimization technique slashes memory costs
#112Wonder how this compares with Microsoft's HeadKV paper [1] which claims a 98% percent reduction in memory while retaining 97% of the performance. [1] https://arxiv.org/html/2410.19258v3
Re: New LLM optimization technique slashes memory costs
#113Re: New LLM optimization technique slashes memory costs
#114Earlier quoted context omitted.
They’re both exploring the same space of optimizing the memory needed by the KV cache which is essentially another name for the context window (no one elides the KV cache as otherwise you’re doing N^2 math to do attention). They’re exploring different approaches to achieve the same goal and they may be both possible to apply simultaneously to reduce the attention mechanism to almost 0 memory usage which would be real…
> they may be both possible to apply simultaneously to reduce the attention mechanism to almost 0 memory usage which would be really cool https://matt.might.net/articles/why-infinite-or-guaranteed-f...
The extent to which using both the techniques together will help will depend on how much overlap there is between the information each ends up discarding.
Re: New LLM optimization technique slashes memory costs
#115Earlier quoted context omitted.
Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.
Microsoft also committed publicly to prioritise security. And Google says they prioritise privacy of their users above all else. I pity the fool that believes anything these corporations put out publicly. Actions matter, words are wind
Re: New LLM optimization technique slashes memory costs
#116Earlier quoted context omitted.
Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.
Google and Microsoft won't do anything that doesn't translate to money. These days are over.
Re: New LLM optimization technique slashes memory costs
#117Earlier quoted context omitted.
Look, I agree that nuclear is difficult, but Google and Microsoft have publicly committed to those projects you’re mentioning. I don’t understand your dismissive tone that all of it is hogwash? This is one of those HN armchair comments.
I feel like taking Google’s commitment to something seriously is one of this things that I can very uncontroversially respond to with “is this your first day?” All but the biggest Google fanboys know that Google is incredibly indecisive and will cut plans at a moment’s notice.
Re: New LLM optimization technique slashes memory costs
#118Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?
Re: New LLM optimization technique slashes memory costs
#119Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?
Re: New LLM optimization technique slashes memory costs
#120Is it possible that after 3-4 years of performance optimizations, both algorithmic and in hardware efficiency, it will turn out that we didn’t really need all of the nuclear plants we’re currently in the process of setting up to satisfy the power demands of AI data centers?