Live data from Hacker News

ChatGPT Explained: A normie's guide to how it works

jonstokes.com

91–100 of 144 posts

Re: ChatGPT Explained: A normie's guide to how it works

#91
post #6

The biggest drawback of LLM is that it never answers with "I don't know" (unless it is some quote) and it just brings bullshit hallucinations which human has to reject as wrong. Thus it is mostly useless for anything serious. Personally I use it to beautify some text, but still have to do a bit of correction to fix b/s or missed context.

> The biggest drawback of LLM is that it never answers with "I don't know" I see this a lot on social media, but it simply isn't true in my experience as someone who uses multiple APIs from OpenAI.

Ask it to “prove how a martingale converges to the zariski measure on a commutative scheme.”

Totally absurd question, totally absurd answer.

Re: ChatGPT Explained: A normie's guide to how it works

#92
post #54

Earlier quoted context omitted.

Not just connect something. For example, at a certain size, GPT models start to "learn" how to do basic arithmetic (addition), even for numbers with multiple digits they've never encountered before. It might look like a small thing, what with computers being quite able to do arithmetic at the base level. But this is a language model, so its a bit different. It learns how to add numbers without "carrying the 1" first,…

Hey! That's a very interesting explanation, could you please provide any references for further/detailed reading on model's abilities to learn to add numbers, with carry and finally with carry across digits?

There is some commentary in the GPT-3 paper https://arxiv.org/pdf/2005.14165.pdf (figure 3.10 and table 3.10)

note: I may have extrapolated a bit more than strictly correct from those two bits, but the accuracy being about 50% indicates carrying is the issue. GPT-3 is at around carrying twice, but hasn't generalized full carrying yet.

Re: ChatGPT Explained: A normie's guide to how it works

#93

Earlier quoted context omitted.

> The biggest drawback of LLM is that it never answers with "I don't know" I see this a lot on social media, but it simply isn't true in my experience as someone who uses multiple APIs from OpenAI.

Ask it to “prove how a martingale converges to the zariski measure on a commutative scheme.” Totally absurd question, totally absurd answer.

I asked ChatGPT Legacy:

"Hello! Can you prove how a martingale converges to the zariski measure on a commutative scheme?"

It replied:

"Hello! The notion of a martingale converging to the Zariski measure on a commutative scheme is a highly advanced and abstract mathematical concept. It is not possible to provide a simple proof for this statement, as it requires a deep understanding of several advanced mathematical concepts such as probability theory, algebraic geometry, and commutative algebra."

It then listed definitions for martingale, Zariski measure, and a commutative scheme, and then concluded with:

"However, to fully understand this concept, you would need to study advanced mathematical topics such as algebraic geometry, commutative algebra, and probability theory."

That seems like a useful enough answer with definitions that effectively converges to "I don't know."

Re: ChatGPT Explained: A normie's guide to how it works

#94
post #73

Earlier quoted context omitted.

Borges, of course, kind of skewered your hopes about 80 years ago.

Bit of context might be nice.

I assume they're referring to Jorge Luis Borges's short story "The Library of Babel," which imagines an enormous library that has a book for every possible sequence of letters and punctuation, all of a certain length (a few hundred pages). The library therefore contains all useful books on any topic in any language (as long as it's covered by the alphabet), but also all useless or inaccurate ones, and of course a vast sea of gibberish. All the knowledge you could want is there, yet unattainable.

Re: ChatGPT Explained: A normie's guide to how it works

#95
post #17

Earlier quoted context omitted.

This seems like a great opportunity for some good old fashioned adversarial training. Post train one LLM to please another LLM, that rates the quality of the first model's responses and calls it on any bullshit! (And vice versa.)

I expect that this will be tried, and I worry about some negative consequences if it works well. It could be a way of generating very effective propaganda, that defeats the efforts of the opposing LLM to call bullshit. On the other hand, diffusion models seem to have replaced GANs for image synthesis, so perhaps there's something I'm missing, or perhaps there's a way to combine both techniques.

I am not sure calling bullshit on images works the same way, given we generally want to invite creativity to image generation.

Although an adversarial extra-finger detector seems in order! Extra fingers are not creativity. They are taboo!!

But for text, there is a lot of structure around bullshitting, and fortunately, the internet is full of examples of people calling bullshit. As long as the bullshit adversary has to give a strong critique to back up any bullshit call, it should work.

And bullshit judgments can be de-bullshitted too.

This would be the same as what our social circles, work colleagues, families, etc. do for us, so a circular firing squad of bullshit judges, who also give gold stars for quality responses, seems like a natural solution for models too.

Re: ChatGPT Explained: A normie's guide to how it works

#96
post #85
post #10

Earlier quoted context omitted.

Disagree. As far as I understand, in this article he argues that in the, say, ChatGPT output, compression happens. But does it really? To make a similarly low resolution metaphor, a “bayesian kaleidoscope” of a language model doesn’t necessarily mean it blurs the “word pixels” it is moving around. Because moving them around, rearranging them is what it essentially does, even if in opaque ways; but not degrading them,…

A better metaphor would be to say it compresses the internet, creates a Markov chain based on that compression. Then to make it work it compresses your prompt so that it can find it in the markov chain, move to the next step, and make a lossy decompression into a text token and adds it. The lossy decompression here is the temperature, higher temperature more lossy and more random words, but since it is lossy in the "…

> creates a Markov chain based on that compression

I dislike that interpretation. It suggests it builds a very basic statistical model, but a very basic statistical model simply wouldn't be able to do what these models can do.

Or alternatively, if you want to consider the model as a markov chain mapping the probability from the previous four thousand tokens to the next token then the space is astronomically large. Beyond astronomically and even economically large, there are ~50,000^4096 possible input states.

Re: ChatGPT Explained: A normie's guide to how it works

#97
I liked the intro of the can opener problem, but I think it's quite funny that given that intro (particularly trying to convince people they're not stupid they just don't know about the weird problem this thing is solving) a large section of the document is about electron orbitals. Possibly the most complicated example of probability distributions many people will know, and many won't know it at all.

> We all learn in basic chemistry that orbitals are just regions of space where an electron is likely to be at any moment

You may be surprised.

Latent space is then introduced in a side note:

> “latent space is the multidimensional space of all likely word sequences the model might output”

So the simple overview is that "Oh hey, it's just like electron orbitals - you know except in a multidimensional space of word sequences"?

The end part is probably the most useful, describing how these things work in a bit more practical sense. Overall this feels like it introduces the fact the model is static and it has a token window in a very complicated way.

Re: ChatGPT Explained: A normie's guide to how it works

#98
post #54
post #30

Earlier quoted context omitted.

So basically, if a pattern that hasn't been encountered before is seen, it will just try to connect "something" together, which is why it does things like predict today's date being in the future etc? The model says, "I don't have a good enough path forwards here, I'll just make one up given the next best thing I have and serve it back"? Maybe this is why Bing is working differently, they've changed the model or the…

Not just connect something. For example, at a certain size, GPT models start to "learn" how to do basic arithmetic (addition), even for numbers with multiple digits they've never encountered before. It might look like a small thing, what with computers being quite able to do arithmetic at the base level. But this is a language model, so its a bit different. It learns how to add numbers without "carrying the 1" first,…

I wonder if someone could make some standardized way to describe/write mathematical proofs - in a very formal language and then train a model that will try to find more proofs for open questions.

Re: ChatGPT Explained: A normie's guide to how it works

#99
post #54
post #30

Earlier quoted context omitted.

So basically, if a pattern that hasn't been encountered before is seen, it will just try to connect "something" together, which is why it does things like predict today's date being in the future etc? The model says, "I don't have a good enough path forwards here, I'll just make one up given the next best thing I have and serve it back"? Maybe this is why Bing is working differently, they've changed the model or the…

Not just connect something. For example, at a certain size, GPT models start to "learn" how to do basic arithmetic (addition), even for numbers with multiple digits they've never encountered before. It might look like a small thing, what with computers being quite able to do arithmetic at the base level. But this is a language model, so its a bit different. It learns how to add numbers without "carrying the 1" first,…

Okay but that's not "learning" how to do basic math in the same way I can't "learn" japanese just by mimicking the mouth noises. Yeah I'll get some of the pronunciation right sometime, maybe even get it in the right order, but only for the listener. To me, it's still just mouth noises.

Re: ChatGPT Explained: A normie's guide to how it works

#100
post #54

Earlier quoted context omitted.

Not just connect something. For example, at a certain size, GPT models start to "learn" how to do basic arithmetic (addition), even for numbers with multiple digits they've never encountered before. It might look like a small thing, what with computers being quite able to do arithmetic at the base level. But this is a language model, so its a bit different. It learns how to add numbers without "carrying the 1" first,…

Okay but that's not "learning" how to do basic math in the same way I can't "learn" japanese just by mimicking the mouth noises. Yeah I'll get some of the pronunciation right sometime, maybe even get it in the right order, but only for the listener . To me, it's still just mouth noises.

If you were given the task "mimic the sound of Japanese as best you can", at first, you would just learn the basic phonology and just try and mimic the general sounds of the language. You would get good at that, and eventually you would be perfectly mimicking the pronunciation of Japanese phonemes. At some point, it would become worth your time to actually learn how different sounds are put together, e.g. Japanese can't have an S sound followed by a T sound. After that, you may start to learn how the different syllables interact. And at some point, you would learn how words are put together to form sentences.

At a certain point, there really is no difference between "being really good at mimicking the sound of Japanese" and actually knowing it, because in order to mimic it to a high level you will have to actually know it. "Mimic the sound of Japanese" is the equivalent task here to "predict the next token in this text".

Post reply on HN