Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

181–190 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#181

Earlier quoted context omitted.

Don’t sleep on Gemini Deep Research feature either. I use it for my car work and it beats ChatGPT’s offering at that price point every time.

I dunno, I use Deep Research from Claude, ChatGPT, and Gemini, and Gemini is the only one that ignores my requests and always produces the most inane high school student wannabe management consultant "report" with introduction and restatement of the problem and background and all that. Its "voice" (the prose, I mean, not text to speech) is so irritating I've stopped using it. The other ones will do the thing I want:…

Gemini is high on hallucination. When I ask it about my own software it not only changes my own name to a similar one common in my language but also makes up stuff about our team saying some stranger works with us (he works in the same niche but that's about it).

It's annoying when it's so confident making up nonsense.

Imo Chat GPT is just a league above when it comes to reliability.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#182

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

It would be interesting to see if that same question against GPT-5 Thinking produces notably better results.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#183
post #169

Earlier quoted context omitted.

This is a common theme with LLMs (and LLM criticism). The context: I was rushing for a train, I ran into Starbucks at the station for a coffee, I noticed they didn't have cake pops and the staff member didn't appear to know what they were. I see three choices here: 1. Since I'm mildly curious about Starbucks and cake pop availability in the UK, I get on the train, open up my laptop and dedicate realistically a solid…

I think what's interesting/telling is you view (3) as less desirable. Alternatively, you could have spent that half hour on the train exercising your own creativity to try and satisfy your curiosity. Whether you're right or wrong doesn't really matter, because as you acknowledge it's not really important enough to you to matter. Picking (2) eliminates all the possible avenues that might have lead you down. I'm not sa…

I think one of my personal core values is that curiosity should never be left unsatiated if ant all possible!

I spent my half hour on the train satiating all sorts of other things instead (like the identity of that curious looking building in Reading).

> Picking (2) eliminates all the possible avenues that might have lead you down.

I don't think that's the case. Using GPT-5 for the Cake Pop question lead me down a bunch of avenues I may never have encountered otherwise - the structure of Starbucks in the UK, the history of their Cake Pops rollout, the fact that checking nutritional and allergy details on their website is a great way to get an "official" list of their products independent of what's on sale in individual stores, and it sparked me to run a separate search for their Cookies and Cream cake pop and find out had been discontinued in the US.

Not bad for typing a couple of prompts on my phone and then spending a few extra minutes with the results after the research task had completed.

Now multiply that by a dozen plus moments of curiosity per day and my intellectual life feels genuinely elevated - I'm being exposed to so many more interesting and varied avenues than if I was manually doing all of the work on a smaller number of curiosities myself.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#184
post #157

I guess the part where I'm still skeptical are: Google is also still pretty good at search (especially if I avoid the AI summary with udm=14). I'll take one of your examples: Britannica to seed Wikipedia. I searched for "wikipedia encyclopedia brtannica". In less than 1 second, I got search results back. I spend maybe 30 seconds scanning the page; past the Wikipedia article on Encyclopedia Britannica, past the Encycl…

I wonder how all this will really change the web. In your manual mode, you a human, are viewing and visiting webpages, but if one never needs to and always interacts with the web through an agent, what does the web need to look like, and will people even bother making websites? Interesting times ahead.

I’ve been thinking about this as well. Instead of making websites, maybe people will make something else, like some future version of MCP tools/servers? E.g. a restaurant could have an “MCP tool” for checking opening hours, reserving a table, etc.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#185

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

I asked gemini to do a deep research on the role of healthcare insurance companies in the decline of general practicioners in the Netherlands. It based its premise mostly on blogs and whitepapers on company websites, who's job it is to sell automation-software.

AI really needs better source-validation. Not just to combat the hallucination of sources (which gemini seems to do 80% of the time), but also to combat low quality sources that happen to correlate well to the question in the prompt.

It's similar to Google having to fight SEO spam blogs, they now need to do the same in the output of their models.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#186

Earlier quoted context omitted.

I dunno, I use Deep Research from Claude, ChatGPT, and Gemini, and Gemini is the only one that ignores my requests and always produces the most inane high school student wannabe management consultant "report" with introduction and restatement of the problem and background and all that. Its "voice" (the prose, I mean, not text to speech) is so irritating I've stopped using it. The other ones will do the thing I want:…

Gemini is high on hallucination. When I ask it about my own software it not only changes my own name to a similar one common in my language but also makes up stuff about our team saying some stranger works with us (he works in the same niche but that's about it). It's annoying when it's so confident making up nonsense. Imo Chat GPT is just a league above when it comes to reliability.

>Imo Chat GPT is just a league above when it comes to reliability.

Which is in my option, the #1 metric an LLM should strive for. It can take quite some time to get anything out of an LLM. If the model turns out to be unreliable/untrustworthy, the value of its output is lost.

It's weird that modern society (in general) so blindly buys in to all of the marketing speak. AI has a very disruptive effect on society, only because we let it happen.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#187

Earlier quoted context omitted.

To clarify I am not accusing you of that. I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous. For example whether a given policy is Georgist. But often people who believe in a given doctrine will see differences as more important than they objectively are. For example, just to continue with socialism, it's common for socialist beli…

Okay, so let me break it down for you: The Silagi paper makes a factual claim. The Silagi paper claims that there was only one significant tax in the German colony of Kiatschou, a single tax on land. The direct primary sources reveal that this is not the case. There were multiple taxes, most significantly large tariffs. Additionally there were two taxes on land, not one -- a conventional land value tax, and a "land i…

Reminds me of my pet peeve with algorithmic playlists, if I ask Siri or Alexa for bossa nova, all I get is different covers of Girl from Ipanema since that's the most played song on every bossa nova album

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#188
post #175

Like Simon I've started to use camera for random ChatGPT research. For one ChatGPT works fantastically at random bird identification (along with pretty much all other features and likely location) - https://xkcd.com/1425/ There is one big failure mode though - ChatGPT hallucinates middle of simple textual OCR tasks! I will feed ChatGPT a simple computer hardware invoice with 10 items - out comes perfect first few ite…

Yeah, I've been disappointed in GPT-5 for OCR - Gemini 2.5 is much better on that front: https://simonwillison.net/2025/Aug/29/the-perils-of-vibe-cod...

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#190

Earlier quoted context omitted.

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

I asked gemini to do a deep research on the role of healthcare insurance companies in the decline of general practicioners in the Netherlands. It based its premise mostly on blogs and whitepapers on company websites, who's job it is to sell automation-software. AI really needs better source-validation. Not just to combat the hallucination of sources (which gemini seems to do 80% of the time), but also to combat low q…

Better source validation is one of the main reasons I'm excited about GPT-5 Thinking for this. It would be interesting to try your Gemini prompts against that and see how the results compare.
Post reply on HN