Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

181–190 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#181

Earlier quoted context omitted.

> Sure, but isn't that moving the goalposts? Why shouldn't we use LLMs + tools if it works? Personally i do not see it like that at all as one is referring to LLMs specifically while the other is referring to LLMs plus a bunch of other stuff around them. It is like person A claiming that GIF files can be used to play Doom deathmatches, person B responding that, no, a GIF file cannot start a Doom deathmatch, it is fun…

At the end of the day LLM + tools is asking the LLM to create a story with very specific points where "tool calls" are parts of the story, and "tool results" are like characters that provide context. The fact that they can output stories like that, with enough accuracy to make it worthwhile is, IMO, proof that they can "do" whatever we say they can do. They can "do" math by creating a story where a character takes NL…

I think you have that last part backwards, it is not the LLM driving the interaction, it is the program that uses the LLM to generate the instructions that does the actual driving - that is the bit that makes the LLM start doing things. Though that is just splitting hairs.

The original point was about the capabilities LLMs themselves since the context was about the technology itself, not what you can do by making them part of a larger system that combines LLMs (perhaps more than one) with other tools.

Depending on the use case and context this distinction may or may not matter, e.g. if you are trying to sell the entire system, it probably is not any more important how the individual parts of the system work than what libraries you used to make the software.

However it can be important in other contexts, like evaluating the abilities of LLMs themselves.

For example i have written a script on my PC that my window manager calls to grab whatever text i have selected on whatever application i'm running and passes it to a program i've written in llama.cpp to load Mistral Small with a prompt that makes it check for spelling and grammar mistakes which in turn produces some script-readable input that another script displays in a window.

This, in a way, is an entire system. This system helps me find grammar and spelling mistakes in the text i have selected when i'm writing documents where i care about finding such mistakes. However it is not Mistral Small that has the functionality of finding grammar and spelling mistakes in my selected text, it only provides the part that does the text checking, the rest is done by other external non-LLM pieces. An LLM cannot intercept keystrokes in my computer, it cannot grab my selected text nor can create a window on my desktop, it doesn't even understand these concepts. In a way this can be thought as a limitation from the perspective of the end result i want, but i work around it with the other software i have attached to it.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#182
You would think Open AI employees have a pretty good grasp of their model capabilities, but even if you don’t, you probably always want to be on the cautious side for every claim you see on the internet.

This just seems to be the Open AI culture, which for better or worse has helped foster the AI hype environment we are currently in.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#183
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

[deleted]

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#184
post #68

To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not. It was a quote-tweet of this: https:…

Am I correct in thinking this is the 2nd such fumble by a major lab? DeepMind released their “matrix multiplication better than SOTA” paper a few months back, which suggested Gemini had uncovered a new way to optimally multiply two matrices in fewer steps than previously known. Then immediately after their announcement, mathematicians pointed out that their newly discovered SOTA had been in the literature for 30-40 y…

That doesn't match my recollection of the AlphaEvolve release.

Some people just read the "48 multiplications for a 4x4 matrix multiplications" part, and thought they found prior art at that performance or better. But they missed that the supposed prior art had tighter requirements on the contents of the matrix, which meant those algorithms were not usable for implementing a recursive divide and conquer algorithm for much larger matrix multiplications.

Here is a HN poster claiming to be one of the authors rebutting the claim of prior art: https://news.ycombinator.com/item?id=43997136

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#185

Earlier quoted context omitted.

Pretty sure you can fill a room with serious researchers that at the very least will doubt about 2) being solved with LLMs, especially when talking about formal planning with pure LLMs and without a planning framwork. PS: So just we're clear: formal planning in AI making a coding plan in Cursor.

> with pure LLMs and without a planning framwork. Sure, but isn't that moving the goalposts? Why shouldn't we use LLMs + tools if it works? If anything it shows that the early detractors weren't even considering this could work. Yann in particular was skeptical that long-context things can happen in LLMs at all. We now have "agents" that can work a problem for hours, with self context trimming, planning to md files,…

> Why shouldn't we use

So weird that you immediately move the goalposts after accusing somebody of moving the goalposts. Nobody on the planet told you not to use "LLMs + tools if they work." You've moved onto an entirely different discussion with a made-up person.

> All of this just works, today.

Also, it definitely doesn't "just work." It slops around, screws up, reinserts bugs, randomly removes features, ignores instructions, lies, and sometimes you get a lucky result or something close enough that you can fix up. Nothing that should be in production.

Not that they're not very cool and very helpful in a lot of ways. But I've found them more helpful in showing me how they would do something, and getting me so angry that they nerd-snipe me into doing it correctly. I have to admit, 1) however, that sometimes I'm not sure that I'd have gotten there if I hadn't seen it not getting there, and 2) sometimes "doing it correctly" involves dumping the context and telling it almost exactly how I want something implemented.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#186

Earlier quoted context omitted.

> because the risk of the models being abused or trafficked is virtually zero. That's not really true. Look at one if the more common uses for AI porn: taking a photo of someone and making them nude. Deepfake porn exists and it does harm

The harms associated with someone creating a deep fake of you are real but they're pretty insignificant compared to the harms associated with being sex trafficked or being exposed to an STI or being unable to find traditional employment after working in the industry.

You couldn’t just photoshop that before ai came out?

What if you get a model that is 99% similar to your “target” - what we do with that?

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#187

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Pretty much, though Google got bad at these things well before LLMs really came on to the scene, and we can all debate which project manager was responsible and the month and year things took a downward turn, but the IMO obvious catalyst was that "Barely Good Enough" search creates more ad impressions, especially when virtually all of the bad results you are serving are links to sites that also serve Google managed a…

It was a very clear point: when Amit Singhal was kicked out for sexual harassment in the me too era. He was the heart of search quality but he went too far when he was drinking.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#188

Earlier quoted context omitted.

What is "it". Gpt-5 auto? Gpt-5 pro? Deep research? These have wildly different hallucination rates.

If these rates are known it would be great for OpenAI to be open about them so customers can make an informed decision

OpenAI has published a great deal of information about hallucination rates, as have the other major LLM providers.

You can't just give one single global hallucination rate since the rates depend on the different use cases and despite the abundant amount of information available to people on how to pick the appropriate tool for a given task, it seems very few people care to take the time to actually first recognize that these LLMs are tools, and that you do need to learn how to use these tools in order to be productive with them.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#189
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

I was reminded how terribly ChatGPT hallucinates this morning when I used it to look up the 7 point measurement locations for skin fold calipers to estimate body fat.

It correctly described the locations in text, then it offered to provide a diagram.

I said “sure”, and it generated an image saying the chest location is on the neck, and a bunch of other clearly incorrect locations for the other measurement sites.

It’s gotten better. But it’s still bad.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#190

Earlier quoted context omitted.

> Debating about "reasoning" or not is not fruitful, IMO. Thats kind of the whole need isn’t it? Humans can automate simple tasks very effectively and cheaply already. If I ask my pro versions of LLM what the Unicode value of a seahorse is, and it shows a picture of a horse and gives me the Unicode value for a third completely related animal then it’s pretty clear it can’t reason itself out of a wet paper bag.

Sorry perhaps I worded that poorly. I meant debating about if context stuffing is or isn't "reasoning". At the end of the day, whatever RL + long context does to LLMs seems to provide good results. Reasoning or not :)

Well that’s my point and what I think the engineers are screaming at the top of their lungs these days.. that it’s net negative. It makes a really good demo but hasn’t won anything except maybe translating and simple graphics generation.
Post reply on HN