Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

161–170 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#161
post #151
post #98

Earlier quoted context omitted.

It's not the Deep Search or Agent Mode. I select "GPT-5 Thinking" from the model picker and make sure its regular search tool is enabled.

> This is excellent for satisfying curiosity, and occasionally useful for more important endeavors as well. Small nit, Simon: satisfying curiosity is the important endeavor. <3

It feels like the difference between someone painstakingly brushing away eons of dirt and calcification on buried dino bones vs just picking them up off the ground.

In the former, the research feels genuine and in the latter it feels hollow and probably fake.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#162
post #108

Earlier quoted context omitted.

I'm also on Android with the Plus subscription and I also get this. It usually reconnects by itself a few seconds later, but if it doesn't, I've found that you can get to the answer by closing the app and reopening it.

I had the same problem and I figured out how to fix it! For Samsungs, Apps ‐> ChatGPT -> Battery -> Unrestricted completely fixed the issue for me, it continues thinking/outputting in the background now. Should be a similiar setting for other Android distributions. Basically, it wasn't the app's fault, the OS is just halting it in the background to save battery.

Thank you! That fixed it for me on Pixel 8 as well. Would be great if the app suggested this as a fix.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#163

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

> I always ask ChatGPT do directly cite and evaluate sources. And try to get it in the mindset of comparing and contrasting arguments for and against. And I find I must argue against its points to see how it reacts.

Same here. But it often produces broken or bogus links.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#164
post #122

Earlier quoted context omitted.

What's different is that LLMs with search tools used to be terrible - they would run a single search, get back 10 results and summarize those. Often the results were bad, so the answer was bad. GPT-5 Thinking (and o3 before it, but very few people tried o3) does a whole lot better then that. It runs multiple searches, then evaluates the results and runs follow-up searches to try to get to a credible result. This is n…

Like I said I have nothing against the blog post or writing about it, that was by no means meant as a criticism of you. And I agree it's worth writing and talking about. What surprises me is that we're in a forum for technology enthusiasts. FWIW Gemini at least has been pretty good at this since late 2024 IMO. As for where things are now, I just ran a comparison with ChatGPT 5 in thinking mode against Google search's…

Yeah, the new Google AI mode is impressive too. I wrote about that here: https://simonwillison.net/2025/Sep/7/ai-mode/

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#165

Earlier quoted context omitted.

Spending thousands of words to essentially say "ChatGPT's search feature works pretty well now" with mundane examples like finding UK cake pop availability or identifying buildings from train windows. This has been done before by less capable models - it's just a rehash. Should we expect newer models getting worse? The breathless "Research Goblin" framing and detailed play-by-play of basic web searches feels like pad…

Yes this feels very AI like too. A ton of prose for very little substance lol. I skipped half the article to get to the point, went back and re-read and didn't miss much.

I don't use AI to generate writing on my blog.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#166
post #157

I guess the part where I'm still skeptical are: Google is also still pretty good at search (especially if I avoid the AI summary with udm=14). I'll take one of your examples: Britannica to seed Wikipedia. I searched for "wikipedia encyclopedia brtannica". In less than 1 second, I got search results back. I spend maybe 30 seconds scanning the page; past the Wikipedia article on Encyclopedia Britannica, past the Encycl…

I suggest trying that experiment again but picking the hardest of my examples to answer with Google, not the easiest.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#167

Earlier quoted context omitted.

So does anyone have any good examples of it effectively avoiding the blogspam and SEO? Or being fooled by it? How often either way?

I find one thing it doesn't do very well is avoiding marketing articles pushed by a brand itself. e.g. if I search is X better than Y, very likely landing on articles by makers of brand X and Y and not a 3rd party reviewer. When I manually search on Google I can spot marketing articles just by the URL.

Have you tried that with GPT-5 Thinking or is this based on your experience with older versions of ChatGPT + search?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#168
It's certainly better than google/bing search (non-ai). To be honest I had been observing google/bing (via duckduckgo) decline of search capabilities over recent years. I "had been" unless I stopped observing it because it went below any acceptable level. TBH the only thing I can find on them nowadays is products, and sometimes general information. All the technical articles, api links, etc. are unfindable. Among others that's why I'm holding to hackernews recently (which was the best thing I learned from my colleague, and it breaks my information bubbles). So basically I'm usually starting with ddg, then go to google, and if failed, falling back to chatgpt which is very accurate nowadays.

Example query: a keyboard stand with music (notes) stand.

-- Disclaimer--

It might be connected to the web enshittification process which has been undergoing for quite some time already.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#169

> Starbucks in the UK don’t sell cake pops! Do a deep investigative dive ... I used to play games on my computer a lot. Not so much anymore, don't really want to lock myself in a room alone and play games. I have kids and a wife, and it feels isolative. But those days I would, and often the hardware I had was underpowered to be able to experience the game in its full glory. I would often spend hours and hours just ho…

This is a common theme with LLMs (and LLM criticism).

The context: I was rushing for a train, I ran into Starbucks at the station for a coffee, I noticed they didn't have cake pops and the staff member didn't appear to know what they were.

I see three choices here:

1. Since I'm mildly curious about Starbucks and cake pop availability in the UK, I get on the train, open up my laptop and dedicate realistically a solid half hour or more to figuring out what's going on.

2. I fire off a research question at GPT-5 Thinking on my mobile phone.

3. I don't do any research at all and leave my mild curiosity unsaturated.

Realistically, I think the choices are between 2 and 3. I was never going to perform a full research project on this myself.

See also: AI-enhanced development makes me more ambitious with my projects, which I wrote in March 2023 and has aged extremely well. https://simonwillison.net/2023/Mar/27/ai-enhanced-developmen...

I do plenty of deep dive research projects myself into topics both useful and pointless - my blog is full of them!

Now I can take on even more.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#170

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information.

I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

Post reply on HN