Live data from Hacker News

How does GPT obtain its ability? Tracing emergent abilities of language models

yaofu.notion.site

171–180 of 205 posts

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#171
post #170

I have yet to see an output from a big language model that doesn’t just look like P(text|internet). I understand that it’s very easy to ascribe all kinds of qualities to these things, but when the corpus is the Internet, the log likelihood of it sounding like a person is not so different from the corpus sounding like a person. These things are impressive enough without any magical thinking.

Isn't it possible that intelligence is P(words|every sequence of words you've ever heard)?

Possible? Maybe. I believe we know enough now, however, to conclude that it’s astonishingly unlikely.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#172
post #170

I have yet to see an output from a big language model that doesn’t just look like P(text|internet). I understand that it’s very easy to ascribe all kinds of qualities to these things, but when the corpus is the Internet, the log likelihood of it sounding like a person is not so different from the corpus sounding like a person. These things are impressive enough without any magical thinking.

Isn't it possible that intelligence is P(words|every sequence of words you've ever heard)?

Then where would sensible words come from, for the first time, in hominid history?

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#173

Earlier quoted context omitted.

> I have yet to see an output from a big language model that doesn’t just look like P(text|internet) True, but the same can be said of many things; e.g. biology just looks like P(reproduction|environment), the economy just looks like P(profit|markets), etc. There can still be rich structure inside, and useful abstractions to describe them.

Yeah and I hope I didn’t come off like I was trying to knock the technical achievement: it’s remarkable along multiple dissensions: at a minimum technical, infrastructural, mathematical (you don’t throw 25-50k A100s for months at something without running some serious numbers first). It’s possible that I’ve just fallen too far under the influence of Deutsch and Marletto, but as someone who has worked on systems like…

It is common to switch to a conservative mode of thought upon entering one's domain of competence;

The question is, though, if the expertise and intuition developed during, say, running XGboost classifiers at scale in AdTech really of much relevance when thinking about large transformer models trained with a self-supervised objective and RLHF?

If you try to study this in depth, these models can do something the usual "datascience"-tier ones commonly cannot: https://arxiv.org/abs/2205.10343 https://moultano.wordpress.com/2020/10/18/why-deep-learning-...

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#174
post #170

I have yet to see an output from a big language model that doesn’t just look like P(text|internet). I understand that it’s very easy to ascribe all kinds of qualities to these things, but when the corpus is the Internet, the log likelihood of it sounding like a person is not so different from the corpus sounding like a person. These things are impressive enough without any magical thinking.

Isn't it possible that intelligence is P(words|every sequence of words you've ever heard)?

No, you’re conflating System 1 thinking and System 2 thinking.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#175

Earlier quoted context omitted.

The required knowledge is around entity recognition, with "Obama" referring to the 44th POTUS, and not somebody else who happens to have the same surname (and there are multiple of them actually, at least 4 given his family).

This can clearly be guessed from a search as well. Popularity can be well defined, and in the case of Obama, there is clearly one much more popular than the others.

The model still needs to infer from the sentence the entity to look it up. It is also the case that this is a relatively simple example as 'Obama' refers to a single class of entities and there is not a lot of ambiguity around resolution of class, only resolution of specific entity.

Take this sentence:

> When was KitKat released?

I could refer to the sweet, or the Android OS. Vastly different classes, and the model here needs to "decide" to ask for more information to disambiguate the class, and if the class is the sweet, then it needs to disambiguate the taste particular flavour possibly, and even ask the geographic location.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#176

Earlier quoted context omitted.

Yeah and I hope I didn’t come off like I was trying to knock the technical achievement: it’s remarkable along multiple dissensions: at a minimum technical, infrastructural, mathematical (you don’t throw 25-50k A100s for months at something without running some serious numbers first). It’s possible that I’ve just fallen too far under the influence of Deutsch and Marletto, but as someone who has worked on systems like…

It is common to switch to a conservative mode of thought upon entering one's domain of competence; The question is, though, if the expertise and intuition developed during, say, running XGboost classifiers at scale in AdTech really of much relevance when thinking about large transformer models trained with a self-supervised objective and RLHF? If you try to study this in depth, these models can do something the usual…

If you’re curious about how we used ensembled weak learners to form a strong enough learner to kick the shit out of everyone except DoubleClick in AdTech, this is a decent start: https://en.m.wikipedia.org/wiki/Ensemble_learning.

By the standards of AdTech, latent space stuff was pretty unproven when I left the game a few years ago.

Low-rank approximations were in some sense implicit in that, but were not at that time an explicit goal.

If I were doing an AdTech system from scratch I’d most likely reach for the kind of recommender systems that you’re alluding to sooner.

Sounds like you know your stuff :)

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#178

Earlier quoted context omitted.

Whose knowledge is trustworthy? We've somehow come to associate certain institutions or scientific authorities with truth when that is about the furthest from real science: "Have no respect whatsoever for authority; forget who said it and instead look what he starts with, where he ends up, and ask yourself, Is it reasonable?" -Richard P. Feynman "One of the great commandments of science is, "Mistrust arguments from a…

I think part of the issue is that it’s easier to test the limits of or a humans knowledge, and ironically with your quotes I think you’ve supplied evidence that trust is crucial, in that the truest expression of those quotes would be to just deliver the payload and not attach any sort of authority by association to it. You can’t trust it’s answers (to be fair that’s the existing status quo), but you also can’t easily…

Those quotes are not about trust, they are about the rhetorical technique of appeal to authority.

The actual payload though is the mistrust of authority exactly because we are all so susceptible to the logical fallacy of appeal to authority masquerading bullshit as truth.

There is no problem to solve here. ChatGPT should never be an authority on anything.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#179
post #35

Earlier quoted context omitted.

A generative system, be it a neural network or a human, needs a way to test ideas in order to align with reality. If testing is available, then it is possible to advance the state of the art. Ideas are cheap, results matter.

Sure, but that doesn’t seem to square with the topic at hand - “why does an infinite truth and lies machine feel less trustworthy than another human”. It just isn’t a question that needs a high degree of abstraction to respond to.

It is sometimes a lie machine because it lacks grounding in verification. Humans get more grounding than language models but even we are not 100% there - remember the antivax hysteria. The most grounded field is science, but even in scientific papers most things don't replicate. Verification is hard on all levels and requires extensive work. In particle physics all scientists clump together around the CERN accelerator as it is the only source of verification they have (almost, I exaggerate a bit).

It's going to be important to develop AI methods to test and verify, I think unverified model outputs are worthless verbiage. Verification can be based on references, code execution, physical simulations, lab experiments and even language based simulations.

In a few years the situation is going to flip, AI is going to become more reliable than humans. Being tested on millions of cases, it will be more trustworthy than us, no human can be tested to that extent. It's going to be interesting to see how we react to super-valid AI. Our guiding role is going to shrink more and more, we will be the children.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#180

Earlier quoted context omitted.

This can clearly be guessed from a search as well. Popularity can be well defined, and in the case of Obama, there is clearly one much more popular than the others.

The model still needs to infer from the sentence the entity to look it up. It is also the case that this is a relatively simple example as 'Obama' refers to a single class of entities and there is not a lot of ambiguity around resolution of class, only resolution of specific entity. Take this sentence: > When was KitKat released? I could refer to the sweet, or the Android OS. Vastly different classes, and the model h…

Yes, but the amount of knowledge necessary to decide how to make those sorts of decisions is far smaller than the amount of knowledge necessary to answer all such questions.
Post reply on HN