Live data from Hacker News

Some Remarks on Large Language Models

gist.github.com

41–50 of 87 posts

Re: Some Remarks on Large Language Models

#41

While not unsolvable, I think the author is understating this problem a lot: > Also, let's put things in perspective: yes, it is enviromentally costly, but we aren't training that many of them, and the total cost is miniscule compared to all the other energy consumptions we humans do. Part of the reason LLMs aren't that big in the grand scheme of things is because they haven't been good enough and businesses haven't…

Think about aluminum smelting. At some point in the past, only a few researchers could smelt aluminum, and while it used a ton of energy, it was just a few research projects. Then, people realized that aluminum was lighter than steel and could replace it... so suddenly everybody was smelting aluminum. The method to do this involves massive amounts of electricity... but it was fine, because the value of the product (to society) was more than high enough to justify it. Eventually, smelters moved to places where there were natural sources of energy... for example, the Columbia Gorge dam was used to power a massive smelter. Guess where Google put their west coast data center? Right there, because aluminum smelting led to a superfund site and we exported those to growing countries for pollution reasons. So there is lots of "free, carbon-neutral" power from hydro plants.

The interesting details are: the companies with large GPU/TPU fleets are already running them in fairly efficient setups, with high utilization (so you're not blowing carbon emissions on idle machines), and can scale those setups if demand increases. This is not irresonsible. And, the scaleup will only happen if the systems are actually useful.

Basically there are 100 other things I'd focus on trimming environment impact for before LLMs.

Re: Some Remarks on Large Language Models

#42
Interesting post. I find myself moving away from the sort of "compare/contrast with humans" mode and more "let's figure out exactly what this machine _is_" way of thinking.

If we look back at the history of mechanical machines, we see a lot of the same kind of debates happening there that we do around AI today -- comparing them to the abilities of humans or animals, arguing that "sure, this machine can do X, but humans can do Y better..." But over time, we've generally stopped doing that as we've gotten used to mechanical machines. I don't know that I've ever heard anyone compare a wheel to leg, for instance, even though both "do" the same thing, because at this point we take wheels for granted. Wheels are much more efficient at transporting objects across a surface in some circumstances, but no one's going around saying "yeah, but they will never be able to climb stairs as well" because, well, at this point we recognize that's not an actual argument we need to have. We know what wheels do and don't.

These AI machine are a fairly novel type of machine, so we don't yet really understand what arguments make sense to have and which ones are unnecessary. But I like these posts that get more into exactly what an LLM _is_, as I find them helpful in understanding better exactly what kind of machine an LLM is. They're not "intelligent" any more than any other machine is (and historically, people have sometimes ascribed intelligence, even sentience, to simple mechanical machines), but that's not so important. Exactly what we'll end up doing with these machines will be very interesting.

Re: Some Remarks on Large Language Models

#43
post #27

Earlier quoted context omitted.

No of course not, nor did I claim they are.

Similar question to one above, I don't follow why the (positive) ad hominem bolsters or detract from the arguments.

Because the claim of the comment I was responding to was that Yoav doesn't experience bias and consequently dismisses it, so in this case my comment is a direct refutation of the argument.

Re: Some Remarks on Large Language Models

#44
post #30

I would like to see the section on "Common-yet-boring" arguments cleaned up a bit. There is a whole category of "researchers" who just spend their time criticizing LLMs with common-yet-boring arguments (Emily Bender is the best example) such as "they cost a lot to train" (uhhh have you seen how much enterprise spends on cloud for non-LLM stuff? Or seen the power consumption of an aluminum smelting plant? Or calcuated…

It looks pretty good as it stands, I think - to spend too much time on these arguments is to play their game.

Having said that, I would add a note about the whole category of ontological or "nothing but" arguments - saying that an LLM is nothing but a fancy database, search engine, autocomplete or whatever. There's an element of question-begging when these statements are prefaced with "they will never lead to machine understanding because...", and beyond that, the more they are conflated with everyday technology, the more noteworthy their performance appears.

Re: Some Remarks on Large Language Models

#45
post #2

Sometimes I read text like this and really enjoy the deep insights and arguments once I filter out the emotion, attitude, or tone. And I wonder if the core of what they're trying to communicate would be better or more efficiently received if the text was more neutral or positive. E.g. you can be 'bearish' on something and point out 'limitations', or you can say 'this is where I think we are' and 'this is how I think…

When reading as well-constructed an article as this one, I tend to assume its tone pretty accurately reflects the author's position.

Re: Some Remarks on Large Language Models

#46
I cannot understand where the boundary between some of the "common-yet-boring arguments" and "real limitations" is. E.g., the ideas that "You cannot learn anything meaningful based only on form" and "It only connects pieces its seen before according to some statistics" are "boring", but the fact that models have no knowledge of knowledge, or knowledge of time, or any understanding of how texts relate to each other is "real". These are essentially the same things! This is what people may mean when they proffer their "boring critiques", if you press them hard enough. Of course Yoav, being abrest of the field, knows all the details and can talk about the problem in more concrete terms, but "vague" and "boring" are still different things.

I also cannot fathom how models can develop a sense of time, or structured knowledge of the world consisting of discrete objects, even with a large dose of RLHF, if the internal representations are continuous, and layer normalised, and otherwise incapable of arriving at any hard-ish, logic-like rules? All these models seem have deep seated architectural limitations, and they are almost at the limit of the available training data. Being non-vague and positive-minded about this doesn't solve the issue. The models can write polite emails and funny reviews of Persian rags in haiku, but they are deeply unreasonable and 100% unreliable. There is hardly a solid business or social case for this stuff.

Re: Some Remarks on Large Language Models

#47

I cannot understand where the boundary between some of the "common-yet-boring arguments" and "real limitations" is. E.g., the ideas that "You cannot learn anything meaningful based only on form" and "It only connects pieces its seen before according to some statistics" are "boring", but the fact that models have no knowledge of knowledge, or knowledge of time, or any understanding of how texts relate to each other is…

Actually, they struggle even with haikus if you care about proper syllable counts.

Re: Some Remarks on Large Language Models

#48
post #12

Earlier quoted context omitted.

Jews are a minority in Israel?

We're the "majority" by virtue of internationally gerrymandered borders. In the region as a whole? Yes, we're an indigenous minority.

The Boers are an indigenous minority in southern Africa, but in the 80s I wouldn't have used the Boers as an example of people who really understand the experience of bias as a minority.

Re: Some Remarks on Large Language Models

#50
post #31
post #2

Sometimes I read text like this and really enjoy the deep insights and arguments once I filter out the emotion, attitude, or tone. And I wonder if the core of what they're trying to communicate would be better or more efficiently received if the text was more neutral or positive. E.g. you can be 'bearish' on something and point out 'limitations', or you can say 'this is where I think we are' and 'this is how I think…

Maybe you can ask chatGPT for a summary in a less emotional tone ;)

I tried but the original article is too long :/
Post reply on HN