Live data from Hacker News

Some Remarks on Large Language Models

gist.github.com

51–60 of 87 posts

Re: Some Remarks on Large Language Models

#51
GPT-3 is limited, but it has delivered a jolt that demands a general reconsideration of machine vs human intelligence. Has it made you change your mind about anything?

At this point for me, the notion of machine "intelligence" is a more reasonable proposition. However this shift is the result of a reconsideration of the binary proposition of "dumb or intelligent like humans".

First, I propose a possible discriminant for "intelligence" vs "computation" to be the ability of an algorithm to brute force compute a response given the input corpus of the 'AI' under consideration, where the machine has provided a reasonable response.

It also seems reasonable to begin to differentiate 'kinds' of intelligence. On this very planet there are a variety of creatures that exhibit some form of intelligence. And they seem to be distinct kinds. Social insects are arguably intelligent. Crows are discussed frequently on hacker news. Fluffy is not entirely dumb either. But are these all the same 'kind' of intelligence?

Putting cards on the table, at this point it seems eminently possible that we will create some form of mechanical insectoid intelligence. I do not believe insects have any need for 'meaning' - form will do. That distinction also takes the sticky 'what is consciousness?' Q out of the equation.

Re: Some Remarks on Large Language Models

#52

Interesting post. I find myself moving away from the sort of "compare/contrast with humans" mode and more "let's figure out exactly what this machine _is_" way of thinking. If we look back at the history of mechanical machines, we see a lot of the same kind of debates happening there that we do around AI today -- comparing them to the abilities of humans or animals, arguing that "sure, this machine can do X, but huma…

It is a relevant argument though, as a reply to claims of "GPT4 will replace doctors and lawyers and programmers in 6 months".

Re: Some Remarks on Large Language Models

#53
post #30

I would like to see the section on "Common-yet-boring" arguments cleaned up a bit. There is a whole category of "researchers" who just spend their time criticizing LLMs with common-yet-boring arguments (Emily Bender is the best example) such as "they cost a lot to train" (uhhh have you seen how much enterprise spends on cloud for non-LLM stuff? Or seen the power consumption of an aluminum smelting plant? Or calcuated…

Loved that section as well. An addendum I'd include is that many of these arguments are boring as criticisms, but super interesting as research areas. AIs burn energy? Great, let's make efficient architectures. AIs embed bias? Let's get better and measuring and aligning bias. AIs don't cite sources? Most humans don't either, but it sure would make the AI more useful if it did...

(As a PS, I've seen that last one mainly as a refutation for the "LLMs are ready to kill search" meme. In that context it's a very valid objection.)

Re: Some Remarks on Large Language Models

#54
> Another way to say it is that the model is "not grounded". The symbols the model operates on are just symbols, and while they can stand in relation to one another, they do not "ground" to any real-world item.

This is what Math is, abstract syntactic rules. GPTs however seem to struggle in particular at counting, probably because their structure does not have a notion of order. I wonder if future LLMs built for math will basically solve all math (if they will be able to find any proof that is provable or not).

Grounding LLMs to images will be super interesting to see though, because images have order and so much of abstract thinking is spatial/geometric in its base. Perhaps those will be the first true AIs

Re: Some Remarks on Large Language Models

#55
post #30

I would like to see the section on "Common-yet-boring" arguments cleaned up a bit. There is a whole category of "researchers" who just spend their time criticizing LLMs with common-yet-boring arguments (Emily Bender is the best example) such as "they cost a lot to train" (uhhh have you seen how much enterprise spends on cloud for non-LLM stuff? Or seen the power consumption of an aluminum smelting plant? Or calcuated…

[deleted]

Re: Some Remarks on Large Language Models

#56

While not unsolvable, I think the author is understating this problem a lot: > Also, let's put things in perspective: yes, it is enviromentally costly, but we aren't training that many of them, and the total cost is miniscule compared to all the other energy consumptions we humans do. Part of the reason LLMs aren't that big in the grand scheme of things is because they haven't been good enough and businesses haven't…

The models will become much smaller, there are already some papers that show promising results with pruned models

And transformers are not even the final model , who knows what will come next

Re: Some Remarks on Large Language Models

#57
"The models are biased, don't cite their sources, and we have no idea if there may be very negative effects on society by machines that very confidently spew truth/garbage mixtures which are very difficult to fact check"

dumb boring critiques, so what? so boring! we'll "be careful", OK? so just shut up!

Re: Some Remarks on Large Language Models

#58

Earlier quoted context omitted.

We're the "majority" by virtue of internationally gerrymandered borders. In the region as a whole? Yes, we're an indigenous minority.

The Boers are an indigenous minority in southern Africa, but in the 80s I wouldn't have used the Boers as an example of people who really understand the experience of bias as a minority.

>The Boers are an indigenous minority in southern Africa,

No they're not.

Re: Some Remarks on Large Language Models

#59

I found the "grounding" explanation provided by human feedback very insightful: > Why is this significant? At the core the model is still doing language modeling, right? learning to predict the next word, based on text alone? Sure, but here the human annotators inject some level of grounding to the text. Some symbols ("summarize", "translate", "formal") are used in a consistent way together with the concept/task they…

the only word grounded there was the word "summary" . There are so many more which are not possible to be delivered in the same way

Re: Some Remarks on Large Language Models

#60

Interesting post. I find myself moving away from the sort of "compare/contrast with humans" mode and more "let's figure out exactly what this machine _is_" way of thinking. If we look back at the history of mechanical machines, we see a lot of the same kind of debates happening there that we do around AI today -- comparing them to the abilities of humans or animals, arguing that "sure, this machine can do X, but huma…

Conversely, these models open up philosophical questions of "exactly what a human is" beyond language abilities. How much of what we think, do, and perceive comes from the use of language?
Post reply on HN