Humans hallucinating about AI.
"OpenAI Researcher Hallucinates GPT-5 Math Breakthrough" could be a headline from The Onion.
OpenAI researcher announced GPT-5 math breakthrough that never happened
61–70 of 258 posts
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#62Yann LeCun's "Hoisted by their own GPTards" is fantastic.
That seems out of character for him - more like something I'd expect from Elon Musk. What's the context I'm missing?
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#63> Summary (from the article) * OpenAI researchers claimed or suggested that GPT-5 had solved unsolved math problems, but in reality, the model only found known results that were unfamiliar to the operator of erdosproblems.com. * Mathematician Thomas Bloom and Deepmind CEO Demis Hassabis criticized the announcement as misleading, leading the researchers to retract or amend their original claims. * According to mathema…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#64My boss always used to say “our only policy is, don’t be the reason we need to create a new policy”. I suspect OpenAI is going to have some new public communication policies going forward.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#65The sad truth about this incident is that it reveals that OpenAI does not have a serious effort to actually work on unsolved math problems.
I realized they jumped the shark when they announced the pivots to ads and porn. Markets haven’t caught on yet.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#66Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#67Yann LeCun's "Hoisted by their own GPTards" is fantastic.
I might be missing context here, but I'm surprised to see Yann using language that plays on 'retard.' That seems out of character for him - more like something I'd expect from Elon Musk. What's the context I'm missing?
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#68The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not.
It was a quote-tweet of this: https://x.com/MarkSellke/status/1979226538059931886?t=OigN6t..., where the author is saying he's "pushing further on this".
The "this" in question is what this second tweet is in turn quote-tweeting: https://x.com/SebastienBubeck/status/1977181716457701775?t=T... -- where the author says "gpt5-pro is superhuman at literature search: [...] it just solved Erdos Problem #339 (listed as open in the official database erdosproblems.com/forum/thread/3…) by realizing that it had actually been solved 20 years ago"
So, reading the thread in order, you get
* SebastienBubeck: "GPT-5 is really good at literature search, it 'solved' an apparently-open problem by finding an existing solution"
* MarkSellke: "Now it's done ten more"
* kevinweil: "Look at this cool stuff we've done!"
I think the problem here is the way quote-tweets work -- you only see the quoted post and not anything that it in turn is quoting. Kevin Weil had the two previous quotes in his context when he did his post and didn't consider the fact that readers would only see the first level, so wouldn't have Sebastien Bubek's post in mind when they read his.That seems like an easy mistake to entirely honestly make, and I think the pile-on is a little unfair.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#69Earlier quoted context omitted.
How so? I wouldn't put much stock into a roque employee announcing something wrong.
That's not any employee, its their VP of Science.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#70Earlier quoted context omitted.
A human being informed of a mistake will usually be able to resolve it and learn something in the process, whereas an LLM is more likely to spiral into nonsense
You must know people without egos. Humans are better at correcting their mistakes, but far worse at admitting them. But yes, as an edge case handler humans still have an edge.