Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

101–110 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#101
post #85

Earlier quoted context omitted.

A counter example to this is that I asked it about NovaMin® 5 minutes ago and it essentially told me to not bother and buy whatever toothpaste has >1450 ppm fluoride.

Such is the nature of probabilistic systems. Generally speaking, LLMs read the top N search results on the topic in question and uncritically summarize them in their answer. Emphasis on uncritically , therefore the quality of LLM answers is strongly correlated with the quality of top search results. Relevant blog post: https://housefresh.com/beware-of-the-google-ai-salesman/

This is why I am so excited about the way GPT-5 uses its search tool.

GPT-4o and most other AI-assisted search systems in the past worked how you describe: they took the top 10 search results and answered uncritically based on those. If the results were junk the answer was too.

GPT-5 Thinking doesn't do that. Take a look at the thinking trace examples I linked to - in many of them it runs a few searches, evaluates the results, finds that they're not credible enough to generate an answer and so continues browsing and searching.

That's why many of the answers take 1-2 minutes to return!

I frequently see it dismiss information from social media and prefer to go to a source with a good reputation for fact-checking (like a credible newspaper) instead.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#102
post #83

Yes, "GPT-5 with thinking" is great at search, but it's horrible that it shows "Network connection lost. Attempting to reconnect..." after you switch away from the app for even just a few seconds before coming back. It's going to take a minute, so why do I need to keep looking at it and can't go read some more Wikipedia in the mean time? This is insanely user hostile. Is it just me who encounters this? I'm on Plus pl…

Weird. That doesn't happen to me on iOS - I can post the question, wait just long enough for it to display "Thinking...." and then go and do something else.

It even shows me a push notification at the top of my screen when the search task has finished.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#103
post #83

Yes, "GPT-5 with thinking" is great at search, but it's horrible that it shows "Network connection lost. Attempting to reconnect..." after you switch away from the app for even just a few seconds before coming back. It's going to take a minute, so why do I need to keep looking at it and can't go read some more Wikipedia in the mean time? This is insanely user hostile. Is it just me who encounters this? I'm on Plus pl…

I'm also on Android with the Plus subscription and I also get this. It usually reconnects by itself a few seconds later, but if it doesn't, I've found that you can get to the answer by closing the app and reopening it.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#104

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

Yes, this is very much my experience too.

Switching to GPT5 Thinking helps a little, but it often misses things that it wouldn't when I was using o3 or o1.

As an example, I asked it if there were any incidents involving Botchan in an Onsen. This is a text that is readily available and must have been trained on; in the book, Botchan goes swimming in the onsen, and then is humiliated when the next time he comes back, there is a sign saying "No swimming in the Onsen".

According to GPT5 it gives me this, which is subtly wrong.

> In the novel, when Botchan goes to Dōgo Onsen, he notes the posted rules of the bath. One of them forbids things like: > “No swimming in the bath.” (泳ぐべからず) > “No roughhousing / rowdy behavior.” (無闇に騒ぐべからず) > Botchan finds these signs funny because he’s exactly the sort of hot-headed, restless character who might be tempted to splash around or make noise. He jokes in his narration that it seems as though the rules were written specifically to keep people like him out.

Incidentally, Dogo Onsen still has the "No swimming sign", or it did when I went 10 years ago.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#105
post #71

Is this the “Web Search”, “Deep Research”, or “Agent Mode” feature of ChatGPT? Navigating their feature set is… fun.

I believe this is just the normal mode. In my experience, you don't have to select the web search option to make it search the web. I wonder why they have web search as an option at this point (to force the llm to search?)

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#106
I agree with this completely; ChatGPT search is perfect for most use cases. I find it to be better than OpenAI's deep research in my experience-- it often uses 2-3x the sources, and has a more comprehensive, well-thought-out report. I'm sure there are still cases where deep research is preferable, but I haven't come across those yet.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#107
post #85

Earlier quoted context omitted.

In practice this means that you get the same content farm answer dressed up as a trustworthy answer without even getting the opportunity to exercise better judgement. God help you if you rely on them for questions about branded products, they happily rephrase the company's marketing materials as facts.

A counter example to this is that I asked it about NovaMin® 5 minutes ago and it essentially told me to not bother and buy whatever toothpaste has >1450 ppm fluoride.

A year ago I asked it to do deep research on Biomin F + a comparison to NovaMin & fluoride. It gave a comprehensive answer detailing the benefits of BioMin & NovaMin over regular fluroide.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#108
post #83

Yes, "GPT-5 with thinking" is great at search, but it's horrible that it shows "Network connection lost. Attempting to reconnect..." after you switch away from the app for even just a few seconds before coming back. It's going to take a minute, so why do I need to keep looking at it and can't go read some more Wikipedia in the mean time? This is insanely user hostile. Is it just me who encounters this? I'm on Plus pl…

I'm also on Android with the Plus subscription and I also get this. It usually reconnects by itself a few seconds later, but if it doesn't, I've found that you can get to the answer by closing the app and reopening it.

I had the same problem and I figured out how to fix it! For Samsungs, Apps ‐> ChatGPT -> Battery -> Unrestricted completely fixed the issue for me, it continues thinking/outputting in the background now. Should be a similiar setting for other Android distributions. Basically, it wasn't the app's fault, the OS is just halting it in the background to save battery.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#109

Earlier quoted context omitted.

Don’t sleep on Gemini Deep Research feature either. I use it for my car work and it beats ChatGPT’s offering at that price point every time.

I've found the same, but I also haven't gained much value out of "deep research" products as a whole. When I last tested them with topics I'm familiar with, I found the quality of research to be poor. These tools seem to spend their time searching for as much content as possible, then they dump it all into a report. I get better outcomes by extensively searching for a handful of top quality sources. Most of the time…

This begs the question of what would be required to get an AI chatbot to emulate the process you (and others, including myself) use manually, and whether it's possible purely through different prompting.

Is the fundamental problem that it weights all sources equally so a bunch of non-experts stating the wrong answer will overpower a single expert saying the correct answer?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#110

Earlier quoted context omitted.

I've found the same, but I also haven't gained much value out of "deep research" products as a whole. When I last tested them with topics I'm familiar with, I found the quality of research to be poor. These tools seem to spend their time searching for as much content as possible, then they dump it all into a report. I get better outcomes by extensively searching for a handful of top quality sources. Most of the time…

This begs the question of what would be required to get an AI chatbot to emulate the process you (and others, including myself) use manually, and whether it's possible purely through different prompting. Is the fundamental problem that it weights all sources equally so a bunch of non-experts stating the wrong answer will overpower a single expert saying the correct answer?

This post has some interesting suggestions about that: https://open.substack.com/pub/mikecaulfield/p/is-the-llm-res...
Post reply on HN