Live data from Hacker News

OpenAI researcher announced GPT-5 math breakthrough that never happened

the-decoder.com

191–200 of 258 posts

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#192
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

I wonder whether for a lot of the search & literature review-type use-cases where people are trying to use GPT-5 and similar we'd honestly be much better off with a really powerful semantic search engine? Any time you ask a chatbot to summarize the literature for you or answer your question, there's a risk it will hallucinate and give you an unreliable answer. Using LLM-generated embeddings for documents to retrieve…

[deleted]

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#193
post #127

Wouldn't be surprised if OpenAI employees are being asked to phrase ( market ) things this way. This is not the first time they claimed GPT-5 "solved" something [1] [1] https://x.com/SebastienBubeck/status/1970875019803910478 edit: full text It's becoming increasingly clear that gpt5 can solve MINOR open math problems, those that would require a day/few days of a good PhD student. Ofc it's not a 100% guarantee, eg be…

[deleted]

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#194
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

I wonder whether for a lot of the search & literature review-type use-cases where people are trying to use GPT-5 and similar we'd honestly be much better off with a really powerful semantic search engine? Any time you ask a chatbot to summarize the literature for you or answer your question, there's a risk it will hallucinate and give you an unreliable answer. Using LLM-generated embeddings for documents to retrieve…

Since you specifically were wondering if something like this exist, I feel okay with mentioning my own tool https://keenious.com since I think it might fit your needs.

Basically we are trying to combine the benefits of chat with normal academic search results using semantic search and keyword search. That way you get the benefit of LLMs but you’re actually engaging with sources like a normal search.

Hope it was what you were looking for!

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#196

You would think Open AI employees have a pretty good grasp of their model capabilities, but even if you don’t, you probably always want to be on the cautious side for every claim you see on the internet. This just seems to be the Open AI culture, which for better or worse has helped foster the AI hype environment we are currently in.

"It is difficult to get a man to understand something, when his salary depends upon his not understanding it,"

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#197
post #86

> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…

Saying it isn't useful is a bit of an overstatement. It can search, churn through 500k words in a few minutes, and come back with summaries, answers, and sources for each point. Should you blindly trust the summary? No. Should you verify key claims by clicking through to the source? Yes. Is it still incredibly useful as a search tool and productivity booster? Absolutely.

If I knew what’s in the paper I don’t need the summary but if I don’t know what’s in the paper I cannot possibly judge the accuracy of its summary.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#198

Earlier quoted context omitted.

In my experience doing literature super-deep-dives, it hallucinates sources about 50% of the time. (For higher-level literature surveys, it's maybe 5%.) Of the other 50% that are real, it's often ~evenly split into sources I'm familiar with and sources I'm not. So it's hugely useful in surfacing papers that I may very well never have found otherwise using e.g. Google Scholar. It's particularly useful in finding relev…

So, the exact stuff Google used to be good at.

Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#199

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.

That's nothing! Next gen will use the entire power output of a small nation for a week, to tell you a nice cake recipe.

Re: OpenAI researcher announced GPT-5 math breakthrough that never happened

#200

Earlier quoted context omitted.

So, the exact stuff Google used to be good at.

Another win for big tech: Google has been enshittified to such a point that you can now spin up a machine that consumes 1000x the power to give you a result that has a coin toss odds of being totally made up.

A search query probably uses about 10x more electricity than a matching LLM query. There's enough wiggle-room depending on the assumptions that they might be about even. There is no way search uses 1/1000th of an LLM.
Post reply on HN