Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

91–100 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#91

I've also found it to be good at digging deep on things I'm curious about, but don't care enough to spend a lot of time on. As an example, I wanted to know how much sugar by weight is in a coffee syrup so I could make my own dupe. My searches were drowned out by marketing material, but ChatGPT found a datasheet with the info I wanted. I would've eventually found it too, but that's too much effort for an unimportant t…

Don’t sleep on Gemini Deep Research feature either. I use it for my car work and it beats ChatGPT’s offering at that price point every time.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#92

These answers take a shockingly long time to resolve considering you can put the questions into Brave search and get basically the same answers in seconds.

With the walls of low quality sites optimized for SEO these days? Call me unconvinced

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#93
I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online:

https://www.fortressofdoors.com/researchers-beware-of-chatgp...

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#94
post #85

Earlier quoted context omitted.

In practice this means that you get the same content farm answer dressed up as a trustworthy answer without even getting the opportunity to exercise better judgement. God help you if you rely on them for questions about branded products, they happily rephrase the company's marketing materials as facts.

A counter example to this is that I asked it about NovaMin® 5 minutes ago and it essentially told me to not bother and buy whatever toothpaste has >1450 ppm fluoride.

Such is the nature of probabilistic systems. Generally speaking, LLMs read the top N search results on the topic in question and uncritically summarize them in their answer. Emphasis on uncritically, therefore the quality of LLM answers is strongly correlated with the quality of top search results.

Relevant blog post: https://housefresh.com/beware-of-the-google-ai-salesman/

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#95
post #21

Nice writeup. This may nudge me to start using chatbots more for this type of queries. I usually use Perplexity or Kagi Assistant instead. Simon, what's your opinion on doing the same with other frontier systems (like Claude?), or is there something specific to ChatGPT+GPT5? I also like the name, nicely encodes some peculiarities of tech. Perhaps we should call AI agents "Goblins" instead.

[deleted]

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#96
post #79

I do miss the earlier "heavy" models that had encyclopedic knowledge vs the new "lighter" models that rely on web search. Relying on web search surfaces a shallow layer of knowledge (thanks to SEO and all the other challenges of ranking web results) vs having ingested / memorized basically the entirety of human written knowledge beyond what's typically reachable within the first 10 results of a web search (eg: digiti…

I feel the opposite. Before I can use information from a model's "internal" knowledge I have to engage in independent research to verify that it's not a hallucination.

Having an LLM generate search strings and then summarize the results does that research up front and automatically, I need only click the sources to verify. Kagi Assistant does this really well.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#97
post #85

Earlier quoted context omitted.

In practice this means that you get the same content farm answer dressed up as a trustworthy answer without even getting the opportunity to exercise better judgement. God help you if you rely on them for questions about branded products, they happily rephrase the company's marketing materials as facts.

A counter example to this is that I asked it about NovaMin® 5 minutes ago and it essentially told me to not bother and buy whatever toothpaste has >1450 ppm fluoride.

What's incredible about that is that you are acting like that was a success story but it is a nuanced topic and it swallowed all the nuance and convinced you.

You're now here telling us how it gave you the right answer, which seems to mostly be due to it confirming your bias.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#98
post #71

Is this the “Web Search”, “Deep Research”, or “Agent Mode” feature of ChatGPT? Navigating their feature set is… fun.

It's not the Deep Search or Agent Mode.

I select "GPT-5 Thinking" from the model picker and make sure its regular search tool is enabled.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#99
post #43
post #18

That post definitely could have been 1/3rd the length

Yeah. I don't understand why the "Official name for the University of Cambridge" example is worth mentioning in the article.

Because it's the simplest example from the last 48 hours of how I've used this tool. I tried to show an illustrative sample of how I am using it.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#100

I've also found it to be good at digging deep on things I'm curious about, but don't care enough to spend a lot of time on. As an example, I wanted to know how much sugar by weight is in a coffee syrup so I could make my own dupe. My searches were drowned out by marketing material, but ChatGPT found a datasheet with the info I wanted. I would've eventually found it too, but that's too much effort for an unimportant t…

Don’t sleep on Gemini Deep Research feature either. I use it for my car work and it beats ChatGPT’s offering at that price point every time.

I've found the same, but I also haven't gained much value out of "deep research" products as a whole. When I last tested them with topics I'm familiar with, I found the quality of research to be poor. These tools seem to spend their time searching for as much content as possible, then they dump it all into a report. I get better outcomes by extensively searching for a handful of top quality sources. Most of the time your question (or at least some subquestions) has already been answered by an expert, and you're better off using their work than sloppily recreating it.
Post reply on HN