Earlier quoted context omitted.
What is "it". Gpt-5 auto? Gpt-5 pro? Deep research? These have wildly different hallucination rates.
I use all of the current versions of ChatGPT, Gemini, and Claude. The hallucination rates are about the same as far as I can tell. It depends mostly on how niche the area is, not which model. They do seem to train on somewhat different sets of academic sources, so it's good to use them all. I'm not talking about deep research or advanced thinking modes -- those are great for some tasks but don't really add anything w…
OpenAI researcher announced GPT-5 math breakthrough that never happened
251–258 of 258 posts
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#252> GPT-5 is proving useful as a literature review assistant No, it does not. It only produces a highly convincing counterfeit. I am honestly happy for people who are satisfied with its output: life is way easier for them than for me. Obviously, the machine discriminates me personally. When I spend hours in the library looking for some engineering-related math made in the 70s-80s, as a last resort measure, I can try to…
I wonder whether for a lot of the search & literature review-type use-cases where people are trying to use GPT-5 and similar we'd honestly be much better off with a really powerful semantic search engine? Any time you ask a chatbot to summarize the literature for you or answer your question, there's a risk it will hallucinate and give you an unreliable answer. Using LLM-generated embeddings for documents to retrieve…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#253Earlier quoted context omitted.
While Yann is clearly brilliant, and has a deeper understanding of the roots of the filed than many of us mortals, I think he's been on a debbie downer trend lately, and more importantly, some of his public stances have been proven wrong in mere months / years after he made them. I remember a public talk, where he was on the stage with some young researcher from MS. (I think it was one of the authors of the "sparks o…
> LLMs can't do math. He went on to "argue" that LLMs trick you with poetry that sounds good, but is highly subjective, and when tested on hard verifiable problems like math, they fail. They really can’t. Token prediction based on context does not reason. You can scramble to submit PRs to ChatGPT to keep up with the “how many Rs in blueberry” kind of problems but it’s clear they can’t even keep up with shitposters on…
Nobody does that. You can't "submit PRs" to an LLM. Although if you pick up new pretraining data you do get people discussing all newly discovered problems, which is a bit of a neat circularity.
> And your 2nd and third point about planning and compounding errors remain challenges.. probably unsolvable with LLM approaches.
Unsolvable in the first place. "Planning" is GOFAI metaphor-based development where they decided humans must do "planning" on no evidence and therefore if they coded something and called it "planning" it would give them intelligence.
Humans don't do or need to do "planning". Much like they don't have or need to have "world models", the other GOFAI obsession.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#254“Mathematician Thomas Bloom, who runs erdosproblems.com, pushed back right away. He called the statements "a dramatic misinterpretation," clarifying that "open" on his site just means he personally doesn't know the solution - not that the problem is actually unsolved.” What mathematician uses this as the definition for “open”? I don’t go around saying that most problems in this textbook are open questions, just becau…
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#255Earlier quoted context omitted.
Pretty much, though Google got bad at these things well before LLMs really came on to the scene, and we can all debate which project manager was responsible and the month and year things took a downward turn, but the IMO obvious catalyst was that "Barely Good Enough" search creates more ad impressions, especially when virtually all of the bad results you are serving are links to sites that also serve Google managed a…
The main reason Google doesn't find good search results anymore is there are no good search results anymore because there are no websites anymore. You can't do it much better.
but the reasons search got hard was that it became profitable to become the "winner" of a search query. It's a hostile market that works to actively undermine you.
AI absolutely will have the same problem if it "takes over" except the websites that win and get your views will not look like blogspam, they will look like (and be) the result of adversarial machine learning.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#256Earlier quoted context omitted.
In my book, chat-based AGI has been reached years ago, when I couldn't reliably distinguish computer from human. Solving problems that humanity couldn't solve is super-AGI or something like that. It's not there indeed.
Beating the Turing Test is not AGI, but it is beating the Turing Test and that was impressive enough when it happened
Which, actually is not a real thing. Nor has it ever really been meaningful.
Trolls on IRC "beat the turing test" with bots that barely even had any functionality.
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#257“Mathematician Thomas Bloom, who runs erdosproblems.com, pushed back right away. He called the statements "a dramatic misinterpretation," clarifying that "open" on his site just means he personally doesn't know the solution - not that the problem is actually unsolved.” What mathematician uses this as the definition for “open”? I don’t go around saying that most problems in this textbook are open questions, just becau…
If a book lists an open problem and you solve it, how would the book know?
Re: OpenAI researcher announced GPT-5 math breakthrough that never happened
#258Earlier quoted context omitted.
Beating the Turing Test is not AGI, but it is beating the Turing Test and that was impressive enough when it happened
So you were impressed by ELIZA right? Because that's what first "beat the turing test" Which, actually is not a real thing. Nor has it ever really been meaningful. Trolls on IRC "beat the turing test" with bots that barely even had any functionality.