Live data from Hacker News

How does GPT obtain its ability? Tracing emergent abilities of language models

yaofu.notion.site

81–90 of 205 posts

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#81
post #21

Not sure about all its filters being RL. Sometimes it seems to flag its output as inappropriate because of a single word (or none at all). Also it has asymmetric behavior, e.g. it will make a joke about men but refuse to make one about women

Competitors will remove the political correctness filters and become more prominent. A Chinese stable diffusion art site is already gaining traction, no filters.

No filters you say? That sounds terrible. Could you provide a link so that I can definitely avoid it.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#82

Earlier quoted context omitted.

Doesn't seem to unreasonable to me. If I asked you "How old is Obama" and you had a data source which had the ages of every person, you don't need to know the answer from memory. Your reasoning tells you that to find out how old someone is you need to check the external resource and what to do with the info once you get it.

The required knowledge is around entity recognition, with "Obama" referring to the 44th POTUS, and not somebody else who happens to have the same surname (and there are multiple of them actually, at least 4 given his family).

Fair, although I’d think it acceptable if the response to your prompt was “which one? I found 4”

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#83
post #21

Not sure about all its filters being RL. Sometimes it seems to flag its output as inappropriate because of a single word (or none at all). Also it has asymmetric behavior, e.g. it will make a joke about men but refuse to make one about women

> Not sure about all its filters being RL. Sometimes it seems to flag its output as inappropriate because of a single word (or none at all). Also it has asymmetric behavior, e.g. it will make a joke about men but refuse to make one about women

Probably because in today's age, making jokes about men is a-okay, but making jokes against women might be perceived as misogynist. Potential misandry is okay, by comparison.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#84

Earlier quoted context omitted.

Yes. GPT doesn't really deal in meanings. Much like autocomplete, it doesn't know what the end of a sentence will be when it starts it. If it randomly chooses different words at the start of a sentence, it may pretend to have a different belief by the end. In reality it doesn't have beliefs any more than a library does. It implicitly contains beliefs (since it's been trained on them) and it can imitate them, but whic…

I see. The assumption here is that one can simulate intelligence without formalizing the notion of meaning. (And if "meaning" is not defined, then the notion of "truth" is impossible to define either). Is my understanding correct?

Some people hope that training will cause it to represent meanings somehow. How to represent meaning isn't well understood.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#85

Earlier quoted context omitted.

Sometimes the meaning of the response is totally different. [ 4 days ago ] The son of my father, but not my brother. Who is he? If a person is the son of the speaker's father but is not the speaker's brother, then that person is the speaker's nephew. A nephew is the son of a person's sibling, so if the speaker's father has a son who is not the speaker's brother, that person is the speaker's nephew. For example, if th…

Yes. GPT doesn't really deal in meanings. Much like autocomplete, it doesn't know what the end of a sentence will be when it starts it. If it randomly chooses different words at the start of a sentence, it may pretend to have a different belief by the end. In reality it doesn't have beliefs any more than a library does. It implicitly contains beliefs (since it's been trained on them) and it can imitate them, but whic…

So if we asked GPT to write a book, it would be hallucinating a chain of words without sticking to any coherent plot. However, we could use a "multi-resolution" approach even with today's version of GPT: at the top level we ask it to write a brief plot for the entire novel, at the next level we'll use this plot as the context and ask to outlines sub-plots of the 3 books in our novel, at the third level we'll use the overall plot and a book's summary as context to generate brief descriptions of chapters in the book, and so on.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#86
post #85

Earlier quoted context omitted.

Yes. GPT doesn't really deal in meanings. Much like autocomplete, it doesn't know what the end of a sentence will be when it starts it. If it randomly chooses different words at the start of a sentence, it may pretend to have a different belief by the end. In reality it doesn't have beliefs any more than a library does. It implicitly contains beliefs (since it's been trained on them) and it can imitate them, but whic…

So if we asked GPT to write a book, it would be hallucinating a chain of words without sticking to any coherent plot. However, we could use a "multi-resolution" approach even with today's version of GPT: at the top level we ask it to write a brief plot for the entire novel, at the next level we'll use this plot as the context and ask to outlines sub-plots of the 3 books in our novel, at the third level we'll use the…

It’s interesting because in the writing world there’s a spectrum with plotters at one end and pantsers at the other. Plotters work similarly to what you’ve suggested, starting with a plot and working their way down to the actual writing. Pantsers just start writing ‘by the seat of their pants’ and see what emerges. Stephen King is famously in the latter camp. Most people fall somewhere in between, having a rough plot in mind and work out the rest as they go along. Would be interesting to see different AIs take different approaches and see what emerged.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#87
post #85

Earlier quoted context omitted.

Yes. GPT doesn't really deal in meanings. Much like autocomplete, it doesn't know what the end of a sentence will be when it starts it. If it randomly chooses different words at the start of a sentence, it may pretend to have a different belief by the end. In reality it doesn't have beliefs any more than a library does. It implicitly contains beliefs (since it's been trained on them) and it can imitate them, but whic…

So if we asked GPT to write a book, it would be hallucinating a chain of words without sticking to any coherent plot. However, we could use a "multi-resolution" approach even with today's version of GPT: at the top level we ask it to write a brief plot for the entire novel, at the next level we'll use this plot as the context and ask to outlines sub-plots of the 3 books in our novel, at the third level we'll use the…

It's worth a try, but I expect you will still get continuity issues between chapter 1 and later chapters. It's not necessarily coherent even at small scale.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#88

Earlier quoted context omitted.

Sometimes the meaning of the response is totally different. [ 4 days ago ] The son of my father, but not my brother. Who is he? If a person is the son of the speaker's father but is not the speaker's brother, then that person is the speaker's nephew. A nephew is the son of a person's sibling, so if the speaker's father has a son who is not the speaker's brother, that person is the speaker's nephew. For example, if th…

Yes. GPT doesn't really deal in meanings. Much like autocomplete, it doesn't know what the end of a sentence will be when it starts it. If it randomly chooses different words at the start of a sentence, it may pretend to have a different belief by the end. In reality it doesn't have beliefs any more than a library does. It implicitly contains beliefs (since it's been trained on them) and it can imitate them, but whic…

And yet it is often able to make surprising references to previous text. This is not just a markov chain, and is capable of what the author describes as chain of thought. I think there are deeper relationships encoded in the model that allow it to keep to a consistent narrative for a very long time. Its beliefs may change between queries but do not, generally, within the context of a single conversation.

Re: How does GPT obtain its ability? Tracing emergent abilities of language models

#90
post #2

Amazing insight, particularly section 6. "- The two important but different abilities of GPT-3.5 are *knowledge* and *reasoning*. Generally, it would be ideal if we could *offload the knowledge part to the outside retrieval system and let the language model only focus on reasoning.* This is because: - The model’s internal knowledge is always cut off at a certain time. The model always needs up-to-date knowledge to an…

How is hosting the knowledge in a large cloud database any different than hosting the model itself in the cloud? Why the need to run "reasoning" locally?
Post reply on HN