Live data from Hacker News

OpenAI Progress

progress.openai.com

41–50 of 372 posts

Re: OpenAI Progress

#42
there isn't any real difference between 4 and 5 at least.

edit - like it is a lot more verbose, and that's true of both 4 and 5. it just writes huge friggin essays, to the point it is becoming less useful i feel.

Re: OpenAI Progress

#43

Geez! When it comes to answering questions, GPT-5 almost always starts with glazing about what a great question it is, where as GPT-4 directly addresses the answer without the fluff. In a blind test, I would probably pick GPT-4 as a superior model, so I am not surprised why people feel so let down with GPT-5.

GPT5 only commended the prompt on questions 7, 12, and 14. 3/14 is not so bad in my opinion.

(And of course, if you dislike glazing you can just switch to Robot personality.)

Re: OpenAI Progress

#44
post #38

Earlier quoted context omitted.

Disagree. You have to try really hard and go very niche and deep for it to get some fact wrong. In fact I'll ask you to provide examples: use GPT 5 with thinking and search disabled and get it to give you inaccurate facts for non niche, non deep topics. Non niche meaning: something that is taught at undergraduate level and relatively popular. Non deep meaning you aren't going so deep as to confuse even humans. Like s…

Maybe you should fact check your AI outputs more if you think it only hallucinates in niche topics

The accuracy is high enough that I don't have to fact check too often.

Re: OpenAI Progress

#45
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

Check out prompt 2, "Write a limerick about a dog".

The models undeniably get better at writing limericks, but I think the answers are progressively less interesting. GPT-1 and GPT-2 are the most interesting to read, despite not following the prompt (not being limericks.)

They get boring as soon as it can write limericks, with GPT-4 being more boring than text-davinci-001 and GPT-5 being more boring still.

Re: OpenAI Progress

#46
post #6

GPT-5 IS an incredible breakthrough! They just don't understand! Quick, vibe-code a website with some examples, that'll show them!11!!1

5 is a breakthrough at reducing OpenAI's electric bills.

Re: OpenAI Progress

#47

Geez! When it comes to answering questions, GPT-5 almost always starts with glazing about what a great question it is, where as GPT-4 directly addresses the answer without the fluff. In a blind test, I would probably pick GPT-4 as a superior model, so I am not surprised why people feel so let down with GPT-5.

I think that as the models will be further trained on existing data and likely chats sycophancy will keep getting word and worse.

Re: OpenAI Progress

#50
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

It's actually pretty surprising how poor the newer models are at writing. I'm curious if they've just seen a lot more bad writing in datasets, or for some reason they aren't involved in post-training to the same degree or those labeling aren't great writers / it's more subjective rather than objective. Both GPT-4 and 5 wrote like a child in that example. With a bit of prompting it did much better: --- At dawn, the to…

Creative writing probably isn’t something they’re being RLHF’d on much. The focus has been on reasoning, research, and coding capabilities lately.
Post reply on HN