Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

121–130 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#121

Earlier quoted context omitted.

I feel the opposite. Before I can use information from a model's "internal" knowledge I have to engage in independent research to verify that it's not a hallucination. Having an LLM generate search strings and then summarize the results does that research up front and automatically, I need only click the sources to verify. Kagi Assistant does this really well.

So does anyone have any good examples of it effectively avoiding the blogspam and SEO? Or being fooled by it? How often either way?

Here's a good article about Google AI mode usually managing to spot and avoid social media misinformation but occasionally falling for it: https://open.substack.com/pub/mikecaulfield/p/is-the-llm-res...

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#122

Yeah this is what people are doing with LLMs every day. I don't quite get what is supposed to be different in the blog post. HN is a bit weird because it's got 99 articles about how evil LLMs are and one article that's like "oh hey I asked an LLM questions and got some answers" and people are like "wow amazing". Not that I mind. I assume Simon just wanted to share some cool nerdy stuff and there's nothing wrong with…

What's different is that LLMs with search tools used to be terrible - they would run a single search, get back 10 results and summarize those.

Often the results were bad, so the answer was bad.

GPT-5 Thinking (and o3 before it, but very few people tried o3) does a whole lot better then that. It runs multiple searches, then evaluates the results and runs follow-up searches to try to get to a credible result.

This is new and worth writing about. LLM search doesn't suck any more.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#123

These answers take a shockingly long time to resolve considering you can put the questions into Brave search and get basically the same answers in seconds.

I like Brave but have found their search to be awful. The AI stuff seems decent enough, but the results populated below are just never what I'm looking for.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#124
post #79

I do miss the earlier "heavy" models that had encyclopedic knowledge vs the new "lighter" models that rely on web search. Relying on web search surfaces a shallow layer of knowledge (thanks to SEO and all the other challenges of ranking web results) vs having ingested / memorized basically the entirety of human written knowledge beyond what's typically reachable within the first 10 results of a web search (eg: digiti…

Most real knowledge is stored outside the head, so intelligent agents can't rely solely on what they've remembered. That's why libraries are so fundamental to universities.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#125

Pretty wild! I wonder how much high school teachers and college professors are struggling with the inevitable usage though? "Do deep internet research and thinking to present as much evidence in favor of the idea that JRR Tolkein's Lord of the Rings trilogy was inspired by Mervyn Peake's Gormenghast series." https://chatgpt.com/share/68bcd796-bf8c-800c-ad7a-51387b1e53...

the thing about students who cheat is most of them are (at least in the context of schoolwork) very lazy and don't care if their work is high quality. i would guess waiting multiple minutes for Thinking mode to give thorough results is very unappealing. 4o or 4o-mini was already good enough for their purposes.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#127

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view.

Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#128

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.

Yea this isn't really a chat gpt problem as a source credibility problem no?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#129

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.

What are you talking about? There are verifiable primary sources that ChatGPT was not citing. There are direct primary historical sources that lay out the full budget of the historical German colony in extreme detail, that directly contradict assertions made in the Silagi paper, that’s not a matter of opinion that’s a matter of verifiable fact.

Also what “axe” am I grinding? The findings are specifically inconvenient for my political beliefs, not confirming my priors! My priors would be flattered if Silagi was correct about everything but the primary sources definitively prove he’s exaggerating.

> You published a blog about that opinion, and you want ChatGPT to say you're to accept your view.

False, and I address this multiple times in the piece. I don’t want ChatGPT to mindlessly agree with me, I want it to discover the primary source documents.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#130
post #122

Yeah this is what people are doing with LLMs every day. I don't quite get what is supposed to be different in the blog post. HN is a bit weird because it's got 99 articles about how evil LLMs are and one article that's like "oh hey I asked an LLM questions and got some answers" and people are like "wow amazing". Not that I mind. I assume Simon just wanted to share some cool nerdy stuff and there's nothing wrong with…

What's different is that LLMs with search tools used to be terrible - they would run a single search, get back 10 results and summarize those. Often the results were bad, so the answer was bad. GPT-5 Thinking (and o3 before it, but very few people tried o3) does a whole lot better then that. It runs multiple searches, then evaluates the results and runs follow-up searches to try to get to a credible result. This is n…

Like I said I have nothing against the blog post or writing about it, that was by no means meant as a criticism of you. And I agree it's worth writing and talking about. What surprises me is that we're in a forum for technology enthusiasts.

FWIW Gemini at least has been pretty good at this since late 2024 IMO.

As for where things are now, I just ran a comparison with ChatGPT 5 in thinking mode against Google search's AI mode across a few questions. They performed the same on the searches I tried and returned substantially the same answer except for some minor variation here or there. Google search is maybe an order of magnitude faster. Google obviously has an advantage here which is that it has full access to their search and ranking index.

And of course the ability to make multiple searches and reason about them for been available for months, maybe almost a year, as deep research mode. I guess the novelty now is you can wait a smaller time and get research that's less deep.

Post reply on HN