Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

241–250 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#241

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.

People love gambling.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#242
post #235

Earlier quoted context omitted.

Right, and I would even go a step further and say the context from SebastienBubeck is stretching "solved" past its breaking point by equating literature research with self-bootsrapped problem solving. When it's later characterized as "previously unsolved" it's doubling down on the same equivocation. Don't get me wrong, effectively surfacing unappreciated research is great and extremely valuable. So there's a real thi…

> Don't get me wrong, effectively surfacing unappreciated research is great and extremely valuable. So there's a real thing here but with the wrong headline attached to it. If I said that I solved a problem, but actually I took a solution for an old book, people would call me a liar. If I was prominent person, it would be academic fraud incident. No one would be saying that "I did extremely valuable thing" or "there…

If you said you "solved", yes - if you said "found a solution" however, there's ambiguity to it, which is part of the confusion here.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#244

Earlier quoted context omitted.

While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…

> LLMs can't do math. He went on to "argue" that LLMs trick you with poetry that sounds good, but is highly subjective, and when tested on hard verifiable problems like math, they fail. They really can’t. Token prediction based on context does not reason. You can scramble to submit PRs to ChatGPT to keep up with the “how many Rs in blueberry” kind of problems but it’s clear they can’t even keep up with shitposters on…

> LLMs can't do math.

Ignoring conversations about 'reasoning', at a fundamental level LLMs do not 'do math' in the way that a calculator or a human does math. Sure we can train bigger and bigger models that give you the impression of this but there are proofs out there that with increased task complexity (in this case multi-digit multiplication) eventually the probability of incorrect predictions converges to 1 (https://arxiv.org/abs/2305.18654)

> And your 2nd and third point about planning and compounding errors remain challenges.. probably unsolvable with LLM approaches.

The same issue applies here, really with any complex multi-step problem.

> Again, mere months later the o series of models came out, and basically proved this point moot. Turns out RL + long context mitigate this fairly well. And a year later, we have all SotA models being able to "solve" problems 100k+ tokens deep.

If you go hands on in any decent size codebase with an agent session length and context size become noticeable issues. Again, mathematically error propagation eventually leads to a 100% chance of error. Yann isn't wrong here, we've just kicked the can a little further down the road. What happens at 200k+ tokens? 500k+ tokens? 1M tokens? The underlying issue of a stochastic system isn't addressed.

>While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately

As he should be. Nothing he said was wrong at a fundamental level. The transformer architecture we have now cannot scale with task complexity. Which is fine, by nature it was not designed for such tasks. The problem is that people see these models work on a subset of small scope complex projects and make claims that go against the underlying architecture. If a model is 'solving' complex or planning tasks but then fails to do similar tasks at a higher complexity it's a sign that there is no underlying deterministic process. What is more likely: the model is genuinely 'planning' or 'solving' complex tasks, or that the model has been trained with enough planning and task related examples that it can make a high probability guess?

> So, yeah, I'd take everything any one singular person says with a huge grain of salt. No matter how brilliant said individual is.

If anything, a guy like Yann with a role such as his at a Mag7 company being realistic (bearish if you are a LLM evangelist) about what the transformer architecture can do is a relief. I'm more inclined to listen to him than a guy like Altman who touts LLMs as the future of humanity meanwhile is path to profitability is AI Tik-Tok, sex chatbots, and a third party way to purchase things from Walmart during a recession.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#245

Earlier quoted context omitted.

The harms associated with someone creating a deep fake of you are real but they're pretty insignificant compared to the harms associated with being sex trafficked or being exposed to an STI or being unable to find traditional employment after working in the industry.

You couldn’t just photoshop that before ai came out? What if you get a model that is 99% similar to your “target” - what we do with that?

Think about the change we saw in combat death tolls when things went from flintlock muskets to machine guns, or when battleships gave way to aircraft, and how many people died unnecessarily due to the generals who were slow to update their tactics. Deepfakes are like that because they lower the cost and improve the success rates enough to be transformative and they cause harm which can’t easily be countered. We’re not going to instantly train society to be better at media literacy and the police can’t just ignore reports of sex crimes, so we’re just having to accept that it’s easier to hurt people than it used to be.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#246

Earlier quoted context omitted.

I felt like I was going crazy when people uncritically accepted the original claim from OpenAI. Have people actually used these models?

From what i have seen, people using AI somehow get the mindset that the AI generated result is "godlike" and the "ultimate truth". Its really, really scary and im not that hopeful for what we will see int he next decade. Once i told a coworker that a piece if his code looked rather funky (without doing a more deep CR), and he told me its "proven correct by AI". I was stunned, and asked him if he knows how LLMs genera…

It just drives me nuts when I see people say things like "yeah I asked ChatGPT about this extremely famous open problem, wish me luck!" Like what do you expect to happen exactly with an engine that can't even consistently keep track of what you wrote ten thousand tokens ago?

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#247
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 y…

no, you are incorrect

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#248
post #90

Earlier quoted context omitted.

We're not even solving problems that humanity can solve. There's been several times where I've posed to models a geometry problem that was novel but possible for me to solve on my own, but LLMs have fallen flat on executing them every time. I'm no mathematician, these are not complex problems, but they're well beyond any AI, even when guided. Instead, they're left to me, my trusty whiteboard, and a non-negligible amo…

I'm pretty sure there are billions of people on the Earth unable to solve your geometry problem. That doesn't make them less human. It's not a benchmark. You should think about something almost any human can do, not selected few. That's the bar. Casual conversation is one of the examples that almost any human can do.

Any human could do it, given the training. Humans largely choosing not to specialize in this way doesn't make them less human, nor did I imply that. Humans have the capacity for it, LLMs fall short universally.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#249
post #4

This is just tit-for-tat clickbait. The researcher’s wording was a bit unclear for sure, but far from incorrect.

I disagree. There is no way to interpret "GPT-5 just found solutions to 10 (!) previously unsolved Erdos problems" as saying something other than GPT-5 having solved them. If it just found existing solutions then they obviously weren't "previously unsolved" so the tweet is wrong. He clearly misunderstood the situation and jumped to the conclusion that GPT-5 had actually solved the problems because that's what he want…

But… it found them… and they were previously marked unsolved…

Note that “solved” does not equal “found”

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#250

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Pretty much, though Google got bad at these things well before LLMs really came on to the scene, and we can all debate which project manager was responsible and the month and year things took a downward turn, but the IMO obvious catalyst was that "Barely Good Enough" search creates more ad impressions, especially when virtually all of the bad results you are serving are links to sites that also serve Google managed a…

The main reason Google doesn't find good search results anymore is there are no good search results anymore because there are no websites anymore. You can't do it much better.
Post reply on HN