Live data from Hacker News

“Don Knuth Plays with ChatGPT” but with ChatGPT-4

gist.github.com

131–140 of 141 posts

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#131
post #81
post #72

Earlier quoted context omitted.

> I would rate a person who provides no sentence at all as performing significantly better The logic failure in the above statement is probably worse than the logic failure of not being able to spontaneously compose a phrase with just 5-letter words - and slipping in one or two with a higher word-count. > I suspect most people could pretty quickly come up with something You'd be very surprised then. Most people fail…

>The logic failure in the above statement And which alleged logic failure is that?

The idea that making a mistake but otherwise fulfilling most of the task is worse than failing to perform any part of it.

Especially in the context of "evaluating the performance of something".

Let's expand this a little to make it even more evident: if the task was "make a paragraph of 100 words using only 5 letter words" and an AI couldn't produce anything at all, whereas another came up with a paragraph of 100 words, except a couple of them had 6 or 4 letters, it would make absolutely no sense to rate the first as "better" than the second in performing the task.

As for understanding the task, the latter exhibits an understanding of it (since it produced a paragraph, and most of the words it used filled the criteria, which wouldn't happen if it chose them randomly), it just made a couple of mistakes (the kind of humans could easily make too in such a task). For the former we can't even be sure if it even understood the task at all.

We don't rate humans that way on performing tasks either (if they got it less than perfect it's worse than not doing it at all). Even math tests at the university level consider the approach and any partial results in the right direction, don't just mark it 0 if there's an error, nor give a higher mark to students who didn't produce anything.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#132

What I find amazing about the original exchange was the profound lack of curiosity Knuth demonstrated. Because the model wasn’t flawless in performance he pinned it as a curiosity that was good at grammar and vacuous otherwise and wasn’t interested to hear how it improves. This reminds me of an awful lot of the computing field in this drama as it plays out. People that literally know how implausible any of these feat…

I can't comprehend this comment. Kunths commentary was glowing praise for the AI's thinking ability (and none of the "it's not AI" BS that is so popular), plus a statement that he believes accuracy is more important than raw power, so he wants "you" to work on that. Knuth commented on GPT 4 at the start, and complimented its power and correctness at the end.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#133
post #40

Earlier quoted context omitted.

ChatGPT: You didn't say 5-non-repeat-letters, human, jez

Both the first and last words have repeating letters, so they fail under that interpretation too. There would have to be a bizarre interpretation that consecutive-repeating letters are counted as one, but non-consecutive are counted separately, for its response to be considered correct. An AI aware of how to optimally answer questions put to it would find the least objectionable interpretation when one is a subset of…

ChatGPT: You didn't say I couldn't use many interpretations on the same phrase, human. ;)

Jokes apart, I think it is all about the correct prompt.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#134
post #81

Earlier quoted context omitted.

>The logic failure in the above statement And which alleged logic failure is that?

The idea that making a mistake but otherwise fulfilling most of the task is worse than failing to perform any part of it. Especially in the context of "evaluating the performance of something". Let's expand this a little to make it even more evident: if the task was "make a paragraph of 100 words using only 5 letter words" and an AI couldn't produce anything at all, whereas another came up with a paragraph of 100 wor…

>The idea that making a mistake but otherwise fulfilling most of the task is worse than failing to perform any part of it.

The are many contexts in which correctness is important. In such contexts, an incorrect answer is often worse than an explicit non-answer.

>We don't rate humans that way on performing tasks either (if they got it less than perfect it's worse than not doing it at all). Even math tests at the university

Standardized tests often rate incorrect answers worse than non-answers, though yes a university maths test in particular isn't likely to be that sort of test.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#135

Earlier quoted context omitted.

How would you get the correct number? I just did two Google searches and can't find the correct answer anywhere in the first page of results ("Novel The Haj chapters" and "Novel The Haj chapter list"). Even looking in the "look inside" preview on the Penguin Randomhouse website doesn't help because it apparently doesn't have a table of contents. I'm not surprised ChatGPT doesn't know and to me the only bad thing is t…

So this is great. Asking Bing 'how many chapters are in The Haj by Leon Uris?' produces the answer: According to my sources, there are 11 chapters in “The Haj” by Leon Uris[1] [1] https://cs.stanford.edu/~knuth/chatGPT20.txt Which is amazing, because of course that document actually includes TWO different explanations of how many chapters are in The Haj - chatGPT's: The novel consists of 51 chapters and an epilogue,…

The plug-ins are generally much, much worse than ChatGPT itself I have found. You are just hoping it stumbled on right answer.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#137

Thank you for specifying ChatGPT-4. So many commenters on the web say they used GPT4 without specifying if they're using the ChatGPT version. ChatGPT-4 is specifically aligned for answering questions better than the base GPT4 model.

The official name for the model has always been GPT-4. OpenAI has not used the term ChatGPT-4.

It makes sense to call the foundation model GPT-4, like for the previous GPT versions. The fine-tunings are not where its core capabilities come from. Bing is also "a" GPT-4, just with different fine-tuning.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#138

Earlier quoted context omitted.

Sleepsort is the most elegant & efficient sorting algorithm

Sleepsort just pushes the sorting task to the task scheduler, which uses sone other algorithm to do the sorting.

That's what makes it the best, it automatically improves as OS's implement more efficient algos

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#139
post #6

Interesting both completely whiff on the number of chapters in the Haj.

How would you get the correct number? I just did two Google searches and can't find the correct answer anywhere in the first page of results ("Novel The Haj chapters" and "Novel The Haj chapter list"). Even looking in the "look inside" preview on the Penguin Randomhouse website doesn't help because it apparently doesn't have a table of contents. I'm not surprised ChatGPT doesn't know and to me the only bad thing is t…

> How would you get the correct number?

You could simply check the book. It’s a shame there is not more literary data in ChatGPT training corpus.

Re: “Don Knuth Plays with ChatGPT” but with ChatGPT-4

#140

Earlier quoted context omitted.

So this is great. Asking Bing 'how many chapters are in The Haj by Leon Uris?' produces the answer: According to my sources, there are 11 chapters in “The Haj” by Leon Uris[1] [1] https://cs.stanford.edu/~knuth/chatGPT20.txt Which is amazing, because of course that document actually includes TWO different explanations of how many chapters are in The Haj - chatGPT's: The novel consists of 51 chapters and an epilogue,…

The plug-ins are generally much, much worse than ChatGPT itself I have found. You are just hoping it stumbled on right answer.

Absolutely - you don’t really need a chat agent to google things for you unless it’s way better at googling than you are. And right now it grabs the first couple of results for the First search it thinks of and mindlessly summarizes them - I can do that myself thanks.
Post reply on HN