This only decreases memory cost of input context window, not the memory cost to load and run the models.
New LLM optimization technique slashes memory costs
21–30 of 227 posts
Re: New LLM optimization technique slashes memory costs
#22Re: New LLM optimization technique slashes memory costs
#23Re: New LLM optimization technique slashes memory costs
#24Does this mean us plebs can run LLMs on gimped VRAM Nvidia lower end cards?
Re: New LLM optimization technique slashes memory costs
#25Earlier quoted context omitted.
True. Microsoft's all in, Apple's all in, Nvidia is selling shovels, insurance companies are all in, police & military are all in, education is all in, office management is all in. Who is left to pump line up?
As far as I know, finance is not all in. I see Goldman Sachs doing experiments, for example, but it doesn't feel like they're convinced yet.
Re: New LLM optimization technique slashes memory costs
#26Earlier quoted context omitted.
doesn't training require inference? so i guess it would help there too?
Training doesn't require inference. It uses back-propagation, a different algorithm.
Re: New LLM optimization technique slashes memory costs
#27[flagged]
Re: New LLM optimization technique slashes memory costs
#28[flagged]
True. Microsoft's all in, Apple's all in, Nvidia is selling shovels, insurance companies are all in, police & military are all in, education is all in, office management is all in. Who is left to pump line up?
Re: New LLM optimization technique slashes memory costs
#29[flagged]
Google Trends make it seem like we're out of the exponential growth phase for LLMs-- search interest is possibly plateauing. A decline in search interest outside of academia makes sense. The groups who can get by on APIs don't care so much how the sausage is made and just want to see prices come down. Interested parties have likely already found tools that work for them. There's definitely some academic interest outs…
I am only half kidding.