Live data from Hacker News

“Don Knuth Plays with ChatGPT” but with ChatGPT-4

gist.github.com

121–130 of 141 posts

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#121
post #100

It makes you wonder why Knuth bothered with an outdated ChatGPT version? He couldn't find someone with access to GPT-4?

He wasn't that interested and probably didn't know there were two versions. Eventually someone did give him the GPT-4 version I think.

[deleted]

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#122
I would not be surprised if these questions become some form of canonical test for future language models.

Obviously, being the work of Knuth, they are extraordinarily insightful in peeling back the first layer of the answer and providing insight to the underlying properties of both the model itself, and the dataset on which it was trained. It also tests the ability to compute (not recite) very specific facts (e.g. when the sun will be directly above Japan), so checks if subroutines and ephemerides specific to this type of data exist.

But beyond the obvious technical merit - there is an alluding property to base our tests on those whom we respect. I used a similar - but far less sophisticated - set of questions when first exploring ChatGPT. But nobody will be drawn to Dotan Cohen's language model benchmarks - rightfully so. The name Knuth has such reverence in the field that I forsee this test, and variations on it to prevent rigging, becoming a canonical test of language models.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#123
post #88

The sequence of these two threads is just too perfect. Almost likely someone is trying to make a point.

How so? Don Knuth wrote about his experience with ChatGPT. It was submitted to HN and made it to the front page. Someone saw this and decided to submit the same questions to GPT-4 and posted the results. This seems like a perfectly normal sequence of events.

Knuth even mentioned GPT-4 and lamented not having access to it for the test.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#124
post #8

>> What is the most beautiful algorithm? > Quicksort Algorithm Definitive proof that AI must be stopped. Ranking quicksort as more elegant than heapsort?!

Beauty is in the eye of the beholder. I look no further than bubble sort -- it is simple enough I can recite it straight away should someone wake me up at modnight.

Bubblesort is the bestsort.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#125
post #6

Interesting both completely whiff on the number of chapters in the Haj.

How would you get the correct number? I just did two Google searches and can't find the correct answer anywhere in the first page of results ("Novel The Haj chapters" and "Novel The Haj chapter list"). Even looking in the "look inside" preview on the Penguin Randomhouse website doesn't help because it apparently doesn't have a table of contents. I'm not surprised ChatGPT doesn't know and to me the only bad thing is t…

You can get the chapter counts from here:

http://www.bookrags.com/studyguide-the-haj/chapanal001.html

On the left side if you click on "Chapters Summary and Analysis" it gives a break down of the book into 5 parts with varying chapter counts:

Part 1 Chapters 1-20 Part 2 Chapters 1-16 Part 3 Chapters 1-10 Part 4 Chapters 1-17 Part 5 Chapters 1-14

Giving a total of 20+16+10+17+14 = 77 chapters

OTOH, I tried with Bing/Creative, telling it to use this link, and it still failed. Perhaps because you need to click on the "summary and analysis" section to expand it to show the info. It seems there is room for web retrieval-augmented LLMs like Bing to improve here and be a bit more agentic.

Interestingly Knuth's own answer to the question, has a typo, and refers to the book as having "four" chapters, while then continuing on to give the chapter counts as above for all five chapters! Something to confuse future GPTs when the training set includes this, perhaps!

https://cs.stanford.edu/~knuth/chatGPT20.txt

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#126

Earlier quoted context omitted.

It sounds like you profoundly misunderstand Knuth, and LLMs. I recommend a dose of Mickens: https://www.youtube.com/watch?v=ajGX7odA87k

I don’t know Knuth. I understand LLMs for precisely what they are, how they’re built, the math behind them, the limits of what they’re doing, and I don’t over estimate the illusion. However while I see people over estimating them I think they’re extrapolating the current state to a state where it’s limits are restricted and augmented with other techniques and models that address their short comings. Lack of agency? W…

> I don’t know Knuth.

Well you should before taking unwarranted potshots at the man. He's done more for humanity than you or I ever will, eh?

Anyway, you do sound like you know about LLMs, so apologies for that bit.

> People look at LLMs and shake their head failing to realize it’s a single model and single technique that we haven’t even attempted to augment and fail to realize that it’s even possible to augment and constrain LLM with other techniques to address their non trivial failings.

I doubt Knuth is doing that, rather I think the whole thing is orthogonal to his life's work. FWIW, I would love to know his thoughts after reading the GPT4 version of the answers to his questions, eh?

- - - - - -

> I think they’re extrapolating the current state to a state where it’s limits are restricted and [not] augmented with other techniques and models that address their short comings.

I think you might have dropped a negation in that sentence?

> Lack of agency? We have agent techniques. Lack of consistency with reality? We have information retrieval and semantic inference systems. LLMs bring an unreasonably powerful ability to semantically interpret in a space of ambiguity and approximate enough reasoning and inference to tie together all the pieces we’ve built into an ensemble model that’s so close to AGI that it likely doesn’t matter.

I agree! I've been saying for a few minutes now that we'll connect these LLMs to empirical feedback devices and they'll become scientists. Schmidhuber says his goal is "to create an automatic scientist and then retire.", eh?

(FWIW I think there are serious metaphysical ramifications of the pseudo- vs. real- AGI issue, but this isn't the forum for that.)

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#127

Thank you for specifying ChatGPT-4. So many commenters on the web say they used GPT4 without specifying if they're using the ChatGPT version. ChatGPT-4 is specifically aligned for answering questions better than the base GPT4 model.

The official name for the model has always been GPT-4. OpenAI has not used the term ChatGPT-4.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#128

Earlier quoted context omitted.

How would you get the correct number? I just did two Google searches and can't find the correct answer anywhere in the first page of results ("Novel The Haj chapters" and "Novel The Haj chapter list"). Even looking in the "look inside" preview on the Penguin Randomhouse website doesn't help because it apparently doesn't have a table of contents. I'm not surprised ChatGPT doesn't know and to me the only bad thing is t…

So this is great. Asking Bing 'how many chapters are in The Haj by Leon Uris?' produces the answer: According to my sources, there are 11 chapters in “The Haj” by Leon Uris[1] [1] https://cs.stanford.edu/~knuth/chatGPT20.txt Which is amazing, because of course that document actually includes TWO different explanations of how many chapters are in The Haj - chatGPT's: The novel consists of 51 chapters and an epilogue,…

> And one perfectly reasonable way of interpreting that bit of raw text is that the answer to "How many chapters are in The Haj by Leon Uris?" is "11".

Only if you can write a sonnet that is also a haiku!

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#129
post #5

Earlier quoted context omitted.

Based on my experiments it usually does get it right (18 correct answers out of 20 attempts), and the failures I got were similar to this one: a single six-letter word in an otherwise correct sentence.

Sam and friends must be giggling all the way to the bank: they have a service that 'probably' gives the correct result and paying customers are happy to retry until it gets it right.

What have you ever bought that is always correct?

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#130
post #40

Earlier quoted context omitted.

ChatGPT: You didn't say 5-non-repeat-letters, human, jez

Both the first and last words have repeating letters, so they fail under that interpretation too. There would have to be a bizarre interpretation that consecutive-repeating letters are counted as one, but non-consecutive are counted separately, for its response to be considered correct. An AI aware of how to optimally answer questions put to it would find the least objectionable interpretation when one is a subset of…

ll is a single letter in Spanish.
Post reply on HN