Live data from Hacker News

OpenAI Progress

progress.openai.com

51–60 of 372 posts

Re: OpenAI Progress

#52

Earlier quoted context omitted.

I literally just had ChatGPT create a Python program and it used .ends_with instead of .endswith. This was with ChatGPT 5. I mean it got a generic built in function of one of the most popular languages in the world wrong.

"but using LLMs for answering factual questions" this was about fact checking. Of course I know LLM's are going to hallucinate in coding sometimes.

So it isn’t a “fact” that the built in Python function that tests whether a string ends with a substring is “endswith”?

See

https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect

If you know that a source isn’t to be believed in an area you know about, why would you trust that source in an area you don’t know about?

Another funny anecdote, ChatGPT just got the Gell-Man effect wrong.

https://chatgpt.com/share/68a0b7af-5e40-8010-b1e3-ee9ff3c8cb...

Re: OpenAI Progress

#53

Geez! When it comes to answering questions, GPT-5 almost always starts with glazing about what a great question it is, where as GPT-4 directly addresses the answer without the fluff. In a blind test, I would probably pick GPT-4 as a superior model, so I am not surprised why people feel so let down with GPT-5.

Change to robot mode

Re: OpenAI Progress

#57
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

GPT 4.5 (not shown here) is by far the best at writing.

Re: OpenAI Progress

#58

Gpt1 is wild a dog ! she did n't want to be the one to tell him that , did n't want to lie to him . but she could n't . What did I just read

The GPT-1 responses really leak how much of the training material was literature. Probably all those torrented books.

Re: OpenAI Progress

#59

Earlier quoted context omitted.

"but using LLMs for answering factual questions" this was about fact checking. Of course I know LLM's are going to hallucinate in coding sometimes.

So it isn’t a “fact” that the built in Python function that tests whether a string ends with a substring is “endswith”? See https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect If you know that a source isn’t to be believed in an area you know about, why would you trust that source in an area you don’t know about? Another funny anecdote, ChatGPT just got the Gell-Man effect wrong. https://chatgpt.com/share/68a0b7af…

It got it right with thinking which was the challenge I posed. https://chatgpt.com/share/68a0b897-f8dc-800b-8799-9be2a8ad54...

Re: OpenAI Progress

#60

I thought the response to "what would you say if you could talk to a future AI" would be "how many r in strawberry".

Can we stop with that outdated meme? What model can't answer that effectively?

Effectively yes. Correctly no.

https://claude.ai/share/dda533a3-6976-46fe-b317-5f9ce4121e76

Post reply on HN