Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

201–210 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#201

Earlier quoted context omitted.

In my experience doing literature super-deep-dives, it hallucinates sources about 50% of the time. (For higher-level literature surveys, it's maybe 5%.) Of the other 50% that are real, it's often ~evenly split into sources I'm familiar with and sources I'm not. So it's hugely useful in surfacing papers that I may very well never have found otherwise using e.g. Google Scholar. It's particularly useful in finding relev…

So, the exact stuff Google used to be good at.

Nope. I'm talking about the stuff keywords are no good at, and which Google Scholar doesn't tend to surface because it's just not cited much or it's from a different niche.

The fact that LLM's understand your question semantically, not just with keyword matching, is huge.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#202

Earlier quoted context omitted.

In my experience doing literature super-deep-dives, it hallucinates sources about 50% of the time. (For higher-level literature surveys, it's maybe 5%.) Of the other 50% that are real, it's often ~evenly split into sources I'm familiar with and sources I'm not. So it's hugely useful in surfacing papers that I may very well never have found otherwise using e.g. Google Scholar. It's particularly useful in finding relev…

What is "it". Gpt-5 auto? Gpt-5 pro? Deep research? These have wildly different hallucination rates.

I use all of the current versions of ChatGPT, Gemini, and Claude.

The hallucination rates are about the same as far as I can tell. It depends mostly on how niche the area is, not which model. They do seem to train on somewhat different sets of academic sources, so it's good to use them all.

I'm not talking about deep research or advanced thinking modes -- those are great for some tasks but don't really add anything when you're just looking for all the sources on a subject, as opposed to a research report.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#203
post #153
post #145

Earlier quoted context omitted.

There's this principle, I forget the name, but how everyone when reading the newspaper, when they read on a subject they're familiar with, will instantly spot all the holes, all the errors. And they will ask themselves, how was this even published in the first place? But then they flip to the next page and they read a story on a subject they're not an expert on and they just accept all of it without question. I think…

The Gell-Mann Amnesia effect. And you're absolutely right, it's extremely pronounced in LLM users.

> The Gell-Mann Amnesia effect. And you're absolutely right, it's extremely pronounced in LLM users.

And I guess a lot of LLM-hype critics have the trait to be much less capable of "being able to flip to the next page and read a story on a subject they're not an expert on and they just accept all of it without question".

Because this is an unusual personality trait, these LLM-hype critics get reprimanded all the time by the "mob" that they don't see the great opportunities that LLMs could bring, even though the LLMs may not be perfect.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#205

Earlier quoted context omitted.

Saying it isn't useful is a bit of an overstatement. It can search, churn through 500k words in a few minutes, and come back with summaries, answers, and sources for each point. Should you blindly trust the summary? No. Should you verify key claims by clicking through to the source? Yes. Is it still incredibly useful as a search tool and productivity booster? Absolutely.

I gave it a PDF recently and asked it to help me generate some tables based on the information there in. I thought I'd be saving myself time. I spent easily twice as long as I would have if it I had done it myself. It kept making trivial mistakes, misunderstanding what was in the PDF, hallucinating, etc.

Last summer I used one of the models to help translate a few German wikipedia pages to English, hoping it would make things easier by keeping all the formatting etc. that I'd lose if I copy-pasted mere content via Google Translate.

I did check the translations were correct as part of this — while my German isn't great, it was sufficient for this — and it was fine up until reaching a long table about the timeline of events relevant to the subject, at which point it couldn't help but make stuff up.

Still useful, but when you find the limits of their competence, there's no point attempting to cajole them to go further. They'll save you whatever % of the task in effort, now you have to do all the rest; it's a waste of effort to think either carrot or stick will get them to succeed if they can't do it in the first few tries.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#206

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.

What questions are you asking LLMs where they're wrong 50% of the time?

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#207
post #162

Earlier quoted context omitted.

> GPT-5 is proving useful as a literature review assistant > No, it does not. > It is excellent when just finding something is enough.

I meant that it obviously fits your needs but not mine

> No, it does not. It only produces a highly convincing counterfeit.

How could you say that with high confidence when you admitted it might be useful for others?

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#208
post #153
post #145

Earlier quoted context omitted.

There's this principle, I forget the name, but how everyone when reading the newspaper, when they read on a subject they're familiar with, will instantly spot all the holes, all the errors. And they will ask themselves, how was this even published in the first place? But then they flip to the next page and they read a story on a subject they're not an expert on and they just accept all of it without question. I think…

The Gell-Mann Amnesia effect. And you're absolutely right, it's extremely pronounced in LLM users.

I've thought of the same analogy, but you know, I've never actually seen someone go "It's wrong about the stuff I understand, but I'll trust it anyway on everything I know nothing about". It's either:

(1) people getting caught using it to do their own jobs for them (i.e. they don't even realise it's wrong about the stuff they do understand);

(2) people who see the problems and therefore don't trust them anywhere at all (i.e. no amnesia, quite sensible reaction);

(3) people who see the problems and therefore limit their use to domains where the answers can be verified (I do this).

--

As an aside, I'm a little worried that I keep spotting turns of phrase that I associate with LLMs, for example where you write "you're absolutely right": I have no idea if that's all just us monkeys copying what we see around us (something we absolutely do), or if you're using that phrase deliberately because of the associations.

The only thing I'm confident of is that you're not doing is karma-farming with an LLM, but that's based on your other comments not sounding at all like LLMs so why would you (oh how surprising it was when I was first accused of being an LLM), but eh, dead internet theory feels more and more real…

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#209
post #101
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

> Kevin Weil had the two previous quotes in his context when he did his post and didn't consider the fact that readers would only see the first level, so wouldn't have Sebastien Bubek's post in mind when they read his. No, Weil said he himself misunderstood Sellke's post[1]. Note Weil's wording (10 previously unsolved Erdos problems) vs. Sellke's wording (10 Erdos problems that were listed as open ). [1] https://x.co…

Also, previous comment omitted the part that now-deleted tweet from Bubeck begins with "Science revolution via AI has officially begun...".

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#210
post #90

Earlier quoted context omitted.

In my book, chat-based AGI has been reached years ago, when I couldn't reliably distinguish computer from human. Solving problems that humanity couldn't solve is super-AGI or something like that. It's not there indeed.

We're not even solving problems that humanity can solve. There's been several times where I've posed to models a geometry problem that was novel but possible for me to solve on my own, but LLMs have fallen flat on executing them every time. I'm no mathematician, these are not complex problems, but they're well beyond any AI, even when guided. Instead, they're left to me, my trusty whiteboard, and a non-negligible amo…

I'm pretty sure there are billions of people on the Earth unable to solve your geometry problem. That doesn't make them less human. It's not a benchmark. You should think about something almost any human can do, not selected few. That's the bar. Casual conversation is one of the examples that almost any human can do.
Post reply on HN