Earlier quoted context omitted.
I feel the opposite. Before I can use information from a model's "internal" knowledge I have to engage in independent research to verify that it's not a hallucination. Having an LLM generate search strings and then summarize the results does that research up front and automatically, I need only click the sources to verify. Kagi Assistant does this really well.
So does anyone have any good examples of it effectively avoiding the blogspam and SEO? Or being fooled by it? How often either way?
GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
121–130 of 268 posts
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#122Yeah this is what people are doing with LLMs every day. I don't quite get what is supposed to be different in the blog post. HN is a bit weird because it's got 99 articles about how evil LLMs are and one article that's like "oh hey I asked an LLM questions and got some answers" and people are like "wow amazing". Not that I mind. I assume Simon just wanted to share some cool nerdy stuff and there's nothing wrong with…
Often the results were bad, so the answer was bad.
GPT-5 Thinking (and o3 before it, but very few people tried o3) does a whole lot better then that. It runs multiple searches, then evaluates the results and runs follow-up searches to try to get to a credible result.
This is new and worth writing about. LLM search doesn't suck any more.
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#123These answers take a shockingly long time to resolve considering you can put the questions into Brave search and get basically the same answers in seconds.
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#124I do miss the earlier "heavy" models that had encyclopedic knowledge vs the new "lighter" models that rely on web search. Relying on web search surfaces a shallow layer of knowledge (thanks to SEO and all the other challenges of ranking web results) vs having ingested / memorized basically the entirety of human written knowledge beyond what's typically reachable within the first 10 results of a web search (eg: digiti…
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#125Pretty wild! I wonder how much high school teachers and college professors are struggling with the inevitable usage though? "Do deep internet research and thinking to present as much evidence in favor of the idea that JRR Tolkein's Lord of the Rings trilogy was inspired by Mervyn Peake's Gormenghast series." https://chatgpt.com/share/68bcd796-bf8c-800c-ad7a-51387b1e53...
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#126Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#127I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...
Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#128I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...
This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#129I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...
This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.
Also what “axe” am I grinding? The findings are specifically inconvenient for my political beliefs, not confirming my priors! My priors would be flattered if Silagi was correct about everything but the primary sources definitively prove he’s exaggerating.
> You published a blog about that opinion, and you want ChatGPT to say you're to accept your view.
False, and I address this multiple times in the piece. I don’t want ChatGPT to mindlessly agree with me, I want it to discover the primary source documents.
Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search
#130Yeah this is what people are doing with LLMs every day. I don't quite get what is supposed to be different in the blog post. HN is a bit weird because it's got 99 articles about how evil LLMs are and one article that's like "oh hey I asked an LLM questions and got some answers" and people are like "wow amazing". Not that I mind. I assume Simon just wanted to share some cool nerdy stuff and there's nothing wrong with…
What's different is that LLMs with search tools used to be terrible - they would run a single search, get back 10 results and summarize those. Often the results were bad, so the answer was bad. GPT-5 Thinking (and o3 before it, but very few people tried o3) does a whole lot better then that. It runs multiple searches, then evaluates the results and runs follow-up searches to try to get to a credible result. This is n…
FWIW Gemini at least has been pretty good at this since late 2024 IMO.
As for where things are now, I just ran a comparison with ChatGPT 5 in thinking mode against Google search's AI mode across a few questions. They performed the same on the searches I tried and returned substantially the same answer except for some minor variation here or there. Google search is maybe an order of magnitude faster. Google obviously has an advantage here which is that it has full access to their search and ranking index.
And of course the ability to make multiple searches and reason about them for been available for months, maybe almost a year, as deep research mode. I guess the novelty now is you can wait a smaller time and get research that's less deep.