Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

131–140 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#131
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

So the first guy said "solved [...] by realizing that it had actually been solved 20 years ago", and the second guy said "found solutions to 10 (!) previously unsolved Erdös problems". Previously unsolved. The context doesn't make that true, does it?

Right, and I would even go a step further and say the context from SebastienBubeck is stretching "solved" past its breaking point by equating literature research with self-bootsrapped problem solving. When it's later characterized as "previously unsolved" it's doubling down on the same equivocation.

Don't get me wrong, effectively surfacing unappreciated research is great and extremely valuable. So there's a real thing here but with the wrong headline attached to it.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#132
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 y…

It's an interesting type of fumble too, because it's easy to (mistakenly!) read it as "LLM tries and fails to solve problem but thinks it solved it" when really it's being credited with originality for discovering or reiterating solutions already out there in the literature.

It sounds like the content of the solutions themselves are perfectly fine, so it's unfortunate that the headline will leave the impression that these are just more hallucinations. They're not hallucinations, they're not wrong, they're just wrongly assigned credit for existing work. Which, you know, where have we heard that one before? It's like the stylistic "borrowing" from artists, but in research form.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#133

Earlier quoted context omitted.

The porn pivot makes perfect sense. Porn is already quite fake and unconvincing and none of that matters.

It might not matter as far as profitability is concerned, ethically the second order effects will be very problematic. I am no puritan but the widespread availability of porn has already affected peoples sexual expectations greatly. AI generated porn is going to remove even more guardrails for behavior previously considered deviant, people will view and bring those expectations back to real life.

To perhaps make the same point as you in a different way, I have no issue with "deviancy" but I think it can accelerate the cycle of chasing a sugar high.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#134
post #112

Earlier quoted context omitted.

The porn / sex-chat one is really disappointing. It seems they've given up even pretending that they are trying to do something beneficial for society. This is just a pure society-be-damned money grab.

My hunch is that they don't have a way to stop anything, so they are creating verticals to at least contain porn, medical, higher-ed users.

Ah... The classic "If we don't do it, someone else will"

Tell that to the thousands of 18 year olds who'll be captured by this predatory service and get AI psychosis

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#135
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 y…

Well, it is important that we have some technology to prevent us from going round in circles by reinventing things, such as search.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#136
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

Saying it isn't useful is a bit of an overstatement. It can search, churn through 500k words in a few minutes, and come back with summaries, answers, and sources for each point.

Should you blindly trust the summary? No. Should you verify key claims by clicking through to the source? Yes. Is it still incredibly useful as a search tool and productivity booster? Absolutely.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#137

Earlier quoted context omitted.

Pretty sure you can fill a room with serious researchers that at the very least will doubt about 2) being solved with LLMs, especially when talking about formal planning with pure LLMs and without a planning framwork. PS: So just we're clear: formal planning in AI making a coding plan in Cursor.

> with pure LLMs and without a planning framwork. Sure, but isn't that moving the goalposts? Why shouldn't we use LLMs + tools if it works? If anything it shows that the early detractors weren't even considering this could work. Yann in particular was skeptical that long-context things can happen in LLMs at all. We now have "agents" that can work a problem for hours, with self context trimming, planning to md files,…

> Sure, but isn't that moving the goalposts?

It can be considered as that, sure, but anytime I see Lecun talking about this, he does recognize that you can patch your way around LLMs, the point is that you are going to hit limits eventually anyways. Specific planning benchmarks like Blockworld and the like show that LLMs (with frameworks) hit limits when they're exposed to out-of-distribution problems, and that's a BIG problem.

> We now have "agents" that can work a problem for hours, with self context trimming, planning to md files, editing those plans and so on. All of this just works, today. We used to dream about it a year ago.

I use them everyday but I still woulnd't really let them work for hours in greenfield projects. And we're seeing big vibe coders like Karpathy say the same.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#138

Earlier quoted context omitted.

> LLMs can't do math. He went on to "argue" that LLMs trick you with poetry that sounds good, but is highly subjective, and when tested on hard verifiable problems like math, they fail. They really can’t. Token prediction based on context does not reason. You can scramble to submit PRs to ChatGPT to keep up with the “how many Rs in blueberry” kind of problems but it’s clear they can’t even keep up with shitposters on…

> They really can’t. Token prediction based on context does not reason. Debating about "reasoning" or not is not fruitful, IMO. It's an endless debate that can go anywhere and nowhere in particular. I try to look at results: https://arxiv.org/pdf/2508.15260 Abstract: > Large Language Models (LLMs) have shown great potential in reasoning tasks through test-time scaling methods like self-consistency with majority votin…

> Debating about "reasoning" or not is not fruitful, IMO.

Thats kind of the whole need isn’t it? Humans can automate simple tasks very effectively and cheaply already. If I ask my pro versions of LLM what the Unicode value of a seahorse is, and it shows a picture of a horse and gives me the Unicode value for a third completely related animal then it’s pretty clear it can’t reason itself out of a wet paper bag.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#139
post #52

Earlier quoted context omitted.

I realized they jumped the shark when they announced the pivots to ads and porn. Markets haven’t caught on yet.

They know where the money is.

I think people hugely overestimate how profitable porn (at least "actual" porn) is. Aylo (the owner of Pornhub) makes peanuts compared to Youtube or Disney.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#140
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

If you’re interested in a literature review tool, I built a public one for some friends in grad school that uses hierarchical mixture models to organize bulk searches and citation networks.

Example: https://platform.sturdystatistics.com/deepdive?search_type=e...

Post reply on HN