Live data from Hacker News

Big LLMs weights are a piece of history

antirez.com

11–20 of 222 posts

Re: Big LLMs weights are a piece of history

#12
post #2

That's really what these are: something analogous to JPEG for language, and queryable in natural language. Tangent: I was thinking the other day: these are not AI in the sense that they are not primarily intelligence . I still don't see much evidence of that. What they do give me is superhuman memory. The main thing I use them for is search, research, and a "rubber duck" that talks back, and it's like having an inter…

I've been looking at it as an "instant reddit comment". I can download a 10G or 80G compressed archive that basically contains the useful parts of the internet, and then I all can use it to synthesize something that is about as good and reliable as a really good reddit comment. Which is nifty. But honestly it's an incredible idea to sell that to businesses.

Reddit seems to puppet humans via engagement farming to do what LLMs do in some cases. Posts are prompts, replies are responses.

Of course they vary widely in quality.

Re: Big LLMs weights are a piece of history

#14
People wanting this would be better off using memory architectures, like how the brain does it. For ML, the simplest approach is putting in memory layers with content-addressible schemes. I have a few links on prototypes in this comment:

https://news.ycombinator.com/item?id=42824960

Re: Big LLMs weights are a piece of history

#15
post #2

That's really what these are: something analogous to JPEG for language, and queryable in natural language. Tangent: I was thinking the other day: these are not AI in the sense that they are not primarily intelligence . I still don't see much evidence of that. What they do give me is superhuman memory. The main thing I use them for is search, research, and a "rubber duck" that talks back, and it's like having an inter…

There's a great article recently by Ted Chiang that elaborated on this idea: https://www.newyorker.com/tech/annals-of-technology/chatgpt-...

Re: Big LLMs weights are a piece of history

#16

People wanting this would be better off using memory architectures, like how the brain does it. For ML, the simplest approach is putting in memory layers with content-addressible schemes. I have a few links on prototypes in this comment: https://news.ycombinator.com/item?id=42824960

Animal brains do not separate long term memory and processing - they are one and the same thing - columnar neural assemblies in the cortex that have learnt to recognize repeated patterns, and in turn activate others.

Re: Big LLMs weights are a piece of history

#18

I love the title "Big LLMs" because it means that we are now making a distinction between big LLMs and minute LLMs and maybe medium LLMs. I'd like to propose the we call them "Tall LLMs", "Grande LLMs", and "Venti LLMs" just to be precise.

But of course these are all flavors of "large", so then we have big large language models, medium large language models, etc, which does indeed make the tall/grande/venti names appropriate, or perhaps similar "all large" condom size names (large, huge, gargantuan).

Re: Big LLMs weights are a piece of history

#19

I love the title "Big LLMs" because it means that we are now making a distinction between big LLMs and minute LLMs and maybe medium LLMs. I'd like to propose the we call them "Tall LLMs", "Grande LLMs", and "Venti LLMs" just to be precise.

What does a 20 LLM signify?

Re: Big LLMs weights are a piece of history

#20
post #2

That's really what these are: something analogous to JPEG for language, and queryable in natural language. Tangent: I was thinking the other day: these are not AI in the sense that they are not primarily intelligence . I still don't see much evidence of that. What they do give me is superhuman memory. The main thing I use them for is search, research, and a "rubber duck" that talks back, and it's like having an inter…

If you want to see what this would actually be like:

https://lcamtuf.coredump.cx/lossifizer/

I think a fun experiment could be to see at what setting the average human can no longer decipher the text.

Post reply on HN