Earlier quoted context omitted.
So, the exact stuff Google used to be good at.
Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.
OpenAI researcher announced GPT-5 math breakthrough that never happened
241–250 of 258 posts
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#242Earlier quoted context omitted.
Right, and I would even go a step further and say the context from SebastienBubeck is stretching "solved" past its breaking point by equating literature research with self-bootsrapped problem solving. When it's later characterized as "previously unsolved" it's doubling down on the same equivocation. Don't get me wrong, effectively surfacing unappreciated research is great and extremely valuable. So there's a real thi…
> Don't get me wrong, effectively surfacing unappreciated research is great and extremely valuable. So there's a real thing here but with the wrong headline attached to it. If I said that I solved a problem, but actually I took a solution for an old book, people would call me a liar. If I was prominent person, it would be academic fraud incident. No one would be saying that "I did extremely valuable thing" or "there…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#243Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#244Earlier quoted context omitted.
While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…
> LLMs can't do math. He went on to "argue" that LLMs trick you with poetry that sounds good, but is highly subjective, and when tested on hard verifiable problems like math, they fail. They really can’t. Token prediction based on context does not reason. You can scramble to submit PRs to ChatGPT to keep up with the “how many Rs in blueberry” kind of problems but it’s clear they can’t even keep up with shitposters on…
Ignoring conversations about 'reasoning', at a fundamental level LLMs do not 'do math' in the way that a calculator or a human does math. Sure we can train bigger and bigger models that give you the impression of this but there are proofs out there that with increased task complexity (in this case multi-digit multiplication) eventually the probability of incorrect predictions converges to 1 (https://arxiv.org/abs/2305.18654)
> And your 2nd and third point about planning and compounding errors remain challenges.. probably unsolvable with LLM approaches.
The same issue applies here, really with any complex multi-step problem.
> Again, mere months later the o series of models came out, and basically proved this point moot. Turns out RL + long context mitigate this fairly well. And a year later, we have all SotA models being able to "solve" problems 100k+ tokens deep.
If you go hands on in any decent size codebase with an agent session length and context size become noticeable issues. Again, mathematically error propagation eventually leads to a 100% chance of error. Yann isn't wrong here, we've just kicked the can a little further down the road. What happens at 200k+ tokens? 500k+ tokens? 1M tokens? The underlying issue of a stochastic system isn't addressed.
>While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately
As he should be. Nothing he said was wrong at a fundamental level. The transformer architecture we have now cannot scale with task complexity. Which is fine, by nature it was not designed for such tasks. The problem is that people see these models work on a subset of small scope complex projects and make claims that go against the underlying architecture. If a model is 'solving' complex or planning tasks but then fails to do similar tasks at a higher complexity it's a sign that there is no underlying deterministic process. What is more likely: the model is genuinely 'planning' or 'solving' complex tasks, or that the model has been trained with enough planning and task related examples that it can make a high probability guess?
> So, yeah, I'd take everything any one singular person says with a huge grain of salt. No matter how brilliant said individual is.
If anything, a guy like Yann with a role such as his at a Mag7 company being realistic (bearish if you are a LLM evangelist) about what the transformer architecture can do is a relief. I'm more inclined to listen to him than a guy like Altman who touts LLMs as the future of humanity meanwhile is path to profitability is AI Tik-Tok, sex chatbots, and a third party way to purchase things from Walmart during a recession.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#245Earlier quoted context omitted.
The harms associated with someone creating a deep fake of you are real but they're pretty insignificant compared to the harms associated with being sex trafficked or being exposed to an STI or being unable to find traditional employment after working in the industry.
You couldn’t just photoshop that before ai came out? What if you get a model that is 99% similar to your “target” - what we do with that?
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#246Earlier quoted context omitted.
I felt like I was going crazy when people uncritically accepted the original claim from OpenAI. Have people actually used these models?
From what i have seen, people using AI somehow get the mindset that the AI generated result is "godlike" and the "ultimate truth". Its really, really scary and im not that hopeful for what we will see int he next decade. Once i told a coworker that a piece if his code looked rather funky (without doing a more deep CR), and he told me its "proven correct by AI". I was stunned, and asked him if he knows how LLMs genera…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#247To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…
Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 y…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#248Earlier quoted context omitted.
We're not even solving problems that humanity can solve. There's been several times where I've posed to models a geometry problem that was novel but possible for me to solve on my own, but LLMs have fallen flat on executing them every time. I'm no mathematician, these are not complex problems, but they're well beyond any AI, even when guided. Instead, they're left to me, my trusty whiteboard, and a non-negligible amo…
I'm pretty sure there are billions of people on the Earth unable to solve your geometry problem. That doesn't make them less human. It's not a benchmark. You should think about something almost any human can do, not selected few. That's the bar. Casual conversation is one of the examples that almost any human can do.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#249This is just tit-for-tat clickbait. The researcher’s wording was a bit unclear for sure, but far from incorrect.
I disagree. There is no way to interpret "GPT-5 just found solutions to 10 (!) previously unsolved Erdos problems" as saying something other than GPT-5 having solved them. If it just found existing solutions then they obviously weren't "previously unsolved" so the tweet is wrong. He clearly misunderstood the situation and jumped to the conclusion that GPT-5 had actually solved the problems because that's what he want…
Note that “solved” does not equal “found”
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#250Earlier quoted context omitted.
So, the exact stuff Google used to be good at.
Pretty much, though Google got bad at these things well before LLMs really came on to the scene, and we can all debate which project manager was responsible and the month and year things took a downward turn, but the IMO obvious catalyst was that "Barely Good Enough" search creates more ad impressions, especially when virtually all of the bad results you are serving are links to sites that also serve Google managed a…