Live data from Hacker News

Big LLMs weights are a piece of history

antirez.com

171–180 of 222 posts

Re: Big LLMs weights are a piece of history

#171

Earlier quoted context omitted.

It's funny you say that, but when travelling abroad I wondered how Europeans and Japanese stay sufficiently hydrated.

For healthy adults, thirst is a perfectly adequate guide to hydration needs. Historically normal patterns of drinking - e.g. water with meals and a few cups of tea or coffee in between - are perfectly sufficient unless you're doing hard physical labour or spending long periods of time outdoors in hot weather. The modern American preoccupation with constantly drinking water is a peculiar cultural phenomenon with no sc…

If you are thirsty you are already dehydrated.

Re: Big LLMs weights are a piece of history

#172

Earlier quoted context omitted.

I want to apologize for this joke in advance. It had to be done. We could take a page from Trump’s book and call them “Beautiful” LLMs. Then we’d have “Big Beautiful LLMs” or just “BBLs” for short. Surely that wouldn’t cause any confusion when Googling.

Weirdly enough, the ITU already chose the superlative for the bigliest radio frequency band to be Tremendous: - Extremely Low Frequency (ELF) - Super Low Frequency (SLF) - Ultra Low Frequency (ULF) - Very Low Frequency (VLF) - Low Frequency (LF) - Medium Frequency (MF) - High Frequency (HF) - Very High Frequency (VHF) - Ultra High Frequency (UHF) - Super High Frequency (SHF) - Extremely High Frequency (EHF) - Tremend…

It bothers me that the level below 3 Hz is not given the name "Tremendously low". Now it's not symmetrical. I hope the ITU is happy...

Re: Big LLMs weights are a piece of history

#173
post #165
post #139

Earlier quoted context omitted.

Yikes, maybe we can take a step back? I'm not sure where this is coming from, frankly. One anodyne summary of my comment above would be: > Let's think and communicate more clearly regarding intelligence. Stuart Russell offers a nice definition: an agent's ability to do a defined task. Maybe something about my comment got you riled up? What was it? You wrote: > What you're doing, with this whole "let's make a bullshit…

I just wanted to share an essay I liked. I didn't think you'd pay it much mind. But I can see now that you are a person devoted to science. If you want to know what I believe, I think computers in the 50's were intelligent. I think gpt2 probably qualified as agi if you take the meaning of the acronym literally. At this point we've blown so far past all expectations in terms of intelligence that I've come to agree wit…

> I reacted negatively to the idea earlier that agency should be considered an aspect of intelligence.

In the hopes of clarifying any misunderstandings of what I mean... I said "agent" in Russell's sense -- a system with goals that has sensors and actuators in some environment. This is a common definition in CS and robotics. (I tend to shy away from using the word "agency" because sometimes it brings along meaning I'm not intending. For example, to many, the word "agency" suggests free will combined with the ability to do something with it.)

I recommend Russell to anyone willing to give him a try. I selected part of his writing that explains why his definition is important to his goals. From page 2 of https://people.eecs.berkeley.edu/~russell/papers/aij-cnt.pdf

> My own motivation for studying AI is to create and understand intelligence as a general property of systems, rather than as a specific attribute of humans. I believe this to be an appropriate goal for the field as a whole...

Re: Big LLMs weights are a piece of history

#174
post #165
post #139

Earlier quoted context omitted.

Yikes, maybe we can take a step back? I'm not sure where this is coming from, frankly. One anodyne summary of my comment above would be: > Let's think and communicate more clearly regarding intelligence. Stuart Russell offers a nice definition: an agent's ability to do a defined task. Maybe something about my comment got you riled up? What was it? You wrote: > What you're doing, with this whole "let's make a bullshit…

I just wanted to share an essay I liked. I didn't think you'd pay it much mind. But I can see now that you are a person devoted to science. If you want to know what I believe, I think computers in the 50's were intelligent. I think gpt2 probably qualified as agi if you take the meaning of the acronym literally. At this point we've blown so far past all expectations in terms of intelligence that I've come to agree wit…

To continue my earlier comment... I prefer not to call an LLM "intelligent" much less "outrageously intelligent". Why? The main reason is communication clarity -- and by communication I mean the notion of a sender communicating a meaning to a receiver. Not just symbolic information (a la Shannon), but a faithful representation in the recipient. The phrase "outrageously intelligent" can have many conflicting interpretations in one's audience. Doing so generates more confusion than clarity.

To say my point a different way, intelligence is contextual. I'm not using "contextual" as some sort of vague excuse to avoid getting into the details. I'm not saying that intelligence cannot be quantified at all. Quite the opposite. Intelligence can be quantified fairly well (in the statistical sense) once a person specifies what they are talking about. Like Russell, I'm saying intelligence is multifaceted and depends on the agent (what sensors it has, what actuators it has), the environment, and the goal.

So what language would I use instead? Rather than speaking about "intelligence" as one thing that people understand and agree on, I would point to task- and goal-specific metrics. How well does a particular LLM do on the GRE? The LSAT?

Sooner or later, people will want to generalize over the specifics. This is where statistical reasoning comes in. With enough evaluations, we can start to discuss generalizations in a way that can be backed up with data. For example, might say things like "LLM X demonstrates high competence on text summarization tasks, provided that it has been pretrained on the relevant concepts" or "LLM Y struggles to discuss normative philosophical issues without falling into sycophancy, unless extensive prompt engineering protocols are used".

I think it helps to remember this: if someone asks "Is X intelligent?", one has the option to reframe the question. One can use it as an opportunity to clarify and teach and get into a substantive conversation. The alternative is suboptimal. But alas, some people demand short answers to poorly framed questions. Unfortunately, the answers they get won't help them.

Re: Big LLMs weights are a piece of history

#177

This doesn't make much sense to me. Unattributed heresay has limited historical value, perhaps zero given that the view of the web most of the weights-available models have is Common Crawl which is itself available for preservation.

I suspect the idea is that sometimes breadth wins out over accuracy. Even if it's unsuited as a primary source, this kind of lossy compression of many many documents might help a conscientious historian discover verifiable things through other routes.

Re: Big LLMs weights are a piece of history

#178
post #170

Earlier quoted context omitted.

I'd prefer to see olive sizes get a renaissance. I was always amused by Super Colossal when following my mom around a store as a little kid. From a random web search, it seems the sizes above Large are: Extra Large, Jumbo, Extra Jumbo, Giant, Colossal, Super Colossal, Mammoth, Super Mammoth, Atlas.

Needs more superlatives. “Biggest” < “Extra Biggest” < “Maximum Biggest”. :D

maximum_biggest_final_2

Re: Big LLMs weights are a piece of history

#179
post #170

Earlier quoted context omitted.

I'd prefer to see olive sizes get a renaissance. I was always amused by Super Colossal when following my mom around a store as a little kid. From a random web search, it seems the sizes above Large are: Extra Large, Jumbo, Extra Jumbo, Giant, Colossal, Super Colossal, Mammoth, Super Mammoth, Atlas.

Needs more superlatives. “Biggest” < “Extra Biggest” < “Maximum Biggest”. :D

"Non Plus Ultra"

Followed by another company introducing their "Plus Ultra" model.

Re: Big LLMs weights are a piece of history

#180
post #142

I love the title "Big LLMs" because it means that we are now making a distinction between big LLMs and minute LLMs and maybe medium LLMs. I'd like to propose the we call them "Tall LLMs", "Grande LLMs", and "Venti LLMs" just to be precise.

I've been labeling LLMS as "teensy", "smol", "mid", "biggg", "yuuge". I've been struggling to figure out where to place the lines between them though.

itsy-bitsy teensy 4B to 29B

smol 30B to 59B

mid 60B to 99B

biggg 100B to 299B

yuuge 300B+

Post reply on HN