Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

121–130 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#121

Earlier quoted context omitted.

The porn pivot makes perfect sense. Porn is already quite fake and unconvincing and none of that matters.

Unfortunately, the porn pivot might be their path to "profitability".

Global porn industry revenue is 100B. They won’t take 10% of that. Real humans are already selling themselves pretty cheap or free en masse.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#122
post #113

Earlier quoted context omitted.

While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…

> AIME is saturated (with tool use) [...] But isn't tool use kinda the crux here? Correct me if I'm mistaken, but wasn't the argument back then on whether LLMs could solve maths problems without e.g. writing python to solve? Cause when "Sparks of AGI" came out in March, prompting gpt-3.5-turbo to code solutions to assist solving maths problems over just solving them directly was already established and seemed like th…

AIME was saturated with tool use (i.e. 99%) for SotA models, but pure NL, no tool still perform "unreasonably well" on the task. Not 100% but still within 90%. And with lots of compute it can reach 99% as well, apparently [1] (@512 rollouts, but still)

[1] - https://arxiv.org/pdf/2508.15260

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#123
post #112

Earlier quoted context omitted.

The porn / sex-chat one is really disappointing. It seems they've given up even pretending that they are trying to do something beneficial for society. This is just a pure society-be-damned money grab.

My hunch is that they don't have a way to stop anything, so they are creating verticals to at least contain porn, medical, higher-ed users.

I'm pretty sure that if they didn't deliberately chose to train on sex chat/stories, etc, then the LLM wouldn't be any good at it. The model isn't getting this capability by training on WikiPedia or Reddit.

So, it's not a matter of them not being able to do a good job of preventing the model from doing it, therefore giving up and instead encouraging it to do it (which anyways makes no sense), but rather them having chosen to train the model to do this. OpenAI is targetting porn as one of their profit centers.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#124

Earlier quoted context omitted.

While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…

> LLMs can't do math. He went on to "argue" that LLMs trick you with poetry that sounds good, but is highly subjective, and when tested on hard verifiable problems like math, they fail. They really can’t. Token prediction based on context does not reason. You can scramble to submit PRs to ChatGPT to keep up with the “how many Rs in blueberry” kind of problems but it’s clear they can’t even keep up with shitposters on…

> They really can’t. Token prediction based on context does not reason.

Debating about "reasoning" or not is not fruitful, IMO. It's an endless debate that can go anywhere and nowhere in particular. I try to look at results:

https://arxiv.org/pdf/2508.15260

Abstract:

> Large Language Models (LLMs) have shown great potential in reasoning tasks through test-time scaling methods like self-consistency with majority voting. However, this approach often leads to diminishing returns in accuracy and high computational overhead. To address these challenges, we introduce Deep Think with Confidence (DeepConf), a simple yet powerful method that enhances both reasoning efficiency and performance at test time. DeepConf leverages modelinternal confidence signals to dynamically filter out low-quality reasoning traces during or after generation. It requires no additional model training or hyperparameter tuning and can be seamlessly integrated into existing serving frameworks. We evaluate DeepConf across a variety of reasoning tasks and the latest open-source models, including Qwen 3 and GPT-OSS series. Notably, on challenging benchmarks such as AIME 2025, DeepConf@512 achieves up to 99.9% accuracy and reduces generated tokens by up to 84.7% compared to full parallel thinking.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#125
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

> "GPT-5 is really good at literature search, it 'solved' an apparently-open problem by finding an existing solution" Survivor bias. I can assure you that GPT-5 fucks up even relatively easy searches. I need to have a very good idea how the results looks like and the ability to test it to be able to use any result from GPT-5. If I throw the dice 1000 times and post about it each time that I got a double six. Am I the…

For literature search that might be ok. It doesn't need to replace any other tools, and if 1/10 it surfaces something you wouldn't have found otherwise it could be worth the time on the dud attempts.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#126
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 years, and was almost certainly in Gemini’s training set.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#127
Wouldn't be surprised if OpenAI employees are being asked to phrase (market) things this way. This is not the first time they claimed GPT-5 "solved" something [1]

[1] https://x.com/SebastienBubeck/status/1970875019803910478

edit: full text

It's becoming increasingly clear that gpt5 can solve MINOR open math problems, those that would require a day/few days of a good PhD student. Ofc it's not a 100% guarantee, eg below gpt5 solves 3/5 optimization conjectures. Imo full impact of this has yet to be internalized...

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#128

Earlier quoted context omitted.

While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…

Pretty sure you can fill a room with serious researchers that at the very least will doubt about 2) being solved with LLMs, especially when talking about formal planning with pure LLMs and without a planning framwork. PS: So just we're clear: formal planning in AI making a coding plan in Cursor.

> with pure LLMs and without a planning framwork.

Sure, but isn't that moving the goalposts? Why shouldn't we use LLMs + tools if it works? If anything it shows that the early detractors weren't even considering this could work. Yann in particular was skeptical that long-context things can happen in LLMs at all. We now have "agents" that can work a problem for hours, with self context trimming, planning to md files, editing those plans and so on. All of this just works, today. We used to dream about it a year ago.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#129
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

In my experience doing literature super-deep-dives, it hallucinates sources about 50% of the time. (For higher-level literature surveys, it's maybe 5%.) Of the other 50% that are real, it's often ~evenly split into sources I'm familiar with and sources I'm not. So it's hugely useful in surfacing papers that I may very well never have found otherwise using e.g. Google Scholar. It's particularly useful in finding relev…

What is "it". Gpt-5 auto? Gpt-5 pro? Deep research? These have wildly different hallucination rates.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#130
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

In my experience doing literature super-deep-dives, it hallucinates sources about 50% of the time. (For higher-level literature surveys, it's maybe 5%.) Of the other 50% that are real, it's often ~evenly split into sources I'm familiar with and sources I'm not. So it's hugely useful in surfacing papers that I may very well never have found otherwise using e.g. Google Scholar. It's particularly useful in finding relev…

So, the exact stuff Google used to be good at.
Post reply on HN