Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

221–230 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#221

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

Yeah trying to make well-researched buying decisions for example is really hard because you'll just quite a lot of opinions dominated by marketing material, which aren't well counterbalanced by the sort of angry Reddit posts or YouTube comments I'd often treat as red flags.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#222
post #77

Pretty wild! I wonder how much high school teachers and college professors are struggling with the inevitable usage though? "Do deep internet research and thinking to present as much evidence in favor of the idea that JRR Tolkein's Lord of the Rings trilogy was inspired by Mervyn Peake's Gormenghast series." https://chatgpt.com/share/68bcd796-bf8c-800c-ad7a-51387b1e53...

I was amused that it used the neologism 'steel-man' -- redundantly, too.

I'm a bit confused, how is it redundant here? It's trying to make the best possible argument from one side that seems to be wrong. Instead of taking the argument at face value, it takes the most charitable understanding of it (not requiring that it happened before, but some parts where perhaps inspired during later revisions) and tries to argue that case.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#223
post #102
post #83

Yes, "GPT-5 with thinking" is great at search, but it's horrible that it shows "Network connection lost. Attempting to reconnect..." after you switch away from the app for even just a few seconds before coming back. It's going to take a minute, so why do I need to keep looking at it and can't go read some more Wikipedia in the mean time? This is insanely user hostile. Is it just me who encounters this? I'm on Plus pl…

Weird. That doesn't happen to me on iOS - I can post the question, wait just long enough for it to display "Thinking...." and then go and do something else. It even shows me a push notification at the top of my screen when the search task has finished.

As a counter I've found that the iOS app is insanely unreliable. I've lost chats, it messes up and says there has been no response, connection lost and more. It's been really bad. Often when it's reported no result and failed, I go to the site and everything is fine. If things fail I no longer retry as that's how I've permanently lost history before (which is insane, don't lose my shit), and go and check the website.

Insane ratio of "app quality" to "magic technology". The models are wild (as someone in the AI mix for the last 20 years or so) and the mobile app and codex integrations are hot garbage.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#224
post #157

I guess the part where I'm still skeptical are: Google is also still pretty good at search (especially if I avoid the AI summary with udm=14). I'll take one of your examples: Britannica to seed Wikipedia. I searched for "wikipedia encyclopedia brtannica". In less than 1 second, I got search results back. I spend maybe 30 seconds scanning the page; past the Wikipedia article on Encyclopedia Britannica, past the Encycl…

Yes, Google with udm=14 is much better than "AI". "AI" might work for the trivia-type questions from this article, which most people aren't interested in to begin with.

It fails completely for complex political or investigative questions where there is no clear answer. Reading a single Wikipedia page is usually a better use of one's time:

You don't have to pretend that you are parallelizing work (which is just for show) while waiting three min for the "AI" answer. You practice speed reading and memory retention. You enhance your own semantic network instead of the network owned and controlled by oligopoly members.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#225

Earlier quoted context omitted.

> It usually take a few back and forths to get to whether or not something is actually true This cuts both ways. I have yet to find an opinion or fact I could not make chatgpt agree with as if objectivly true. Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. (Which some already are in 1 way or another)

I'm not sure teaching it in school is actually going to help. Most people will tell you that of course you need to look at primary sources to verify claims - and then turn around and believe the first thing they here from LLM, Redditor, Wiki article, etc. Even worse, many people get openly hostile to the idea that people should verify claims - "what, you don't believe me?"/"everyone here has been telling you this is…

There were actual Wikipedians arguing not to take a wiki with a grain of salt? If I was in that discussion, I must have missed those posts. Can you link an example?

If you mean whether Wikipedia is unreliable? That's a different story, everything is unreliable. Wikipedia just happens to be potentially less unreliable than many (typically) (if used correctly) (#include caveats.h) .

Sources are like power tools. Use them with respect and caution.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#226

Earlier quoted context omitted.

I feel the opposite. Before I can use information from a model's "internal" knowledge I have to engage in independent research to verify that it's not a hallucination. Having an LLM generate search strings and then summarize the results does that research up front and automatically, I need only click the sources to verify. Kagi Assistant does this really well.

So does anyone have any good examples of it effectively avoiding the blogspam and SEO? Or being fooled by it? How often either way?

Bulk search is the only thing where I’ve been consistently impressed with LLMs.

But, like the parent, I’m using the Kagi assistant.

So the answer here might be “search for 5 things and pull the relevant results” works incredibly well, but first you have to build an extremely good search engine that lets the user filter out spam sites.

That said, this isn’t magic, it’s just automated an hour of googling. If the content doesn’t exist you won’t find it.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#227
I just can't get over how gleeful the author sounds in wasting compute on "an often unreasonable amount of work to search the internet and figure out an answer."

Is that the goal? Send this thing off on a wild goose chase, and hope it comes back with the right answer no matter the cost?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#228
post #209

Earlier quoted context omitted.

> It usually take a few back and forths to get to whether or not something is actually true This cuts both ways. I have yet to find an opinion or fact I could not make chatgpt agree with as if objectivly true. Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. (Which some already are in 1 way or another)

> Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. You are very optimistic. Look at all other skills we are trying to teach in school. 'Critical thinking' has been at the top of nearly every curriculum you can point a finger at for quite a while now. To minimal effect. Or just look at how much math we are trying to teach the kids, and what the…

Perhaps a bit optimistic, but this can be shown in real time: the situation, cause, and effect.

Critical thinking is a much more general skill which is applicable anywhere, thus quicker to be 'buried' under other learned behavior.

This skill has an obvious trigger; you're using AI, which means you should be aware of this.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#229

I just can't get over how gleeful the author sounds in wasting compute on "an often unreasonable amount of work to search the internet and figure out an answer." Is that the goal? Send this thing off on a wild goose chase, and hope it comes back with the right answer no matter the cost?

Uh, people have wasted entire lifetimes chasing wild goose. Newton and Einstein both spent the latter halves of their lives :( despite being geniuses.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#230
post #101

Earlier quoted context omitted.

Such is the nature of probabilistic systems. Generally speaking, LLMs read the top N search results on the topic in question and uncritically summarize them in their answer. Emphasis on uncritically , therefore the quality of LLM answers is strongly correlated with the quality of top search results. Relevant blog post: https://housefresh.com/beware-of-the-google-ai-salesman/

This is why I am so excited about the way GPT-5 uses its search tool. GPT-4o and most other AI-assisted search systems in the past worked how you describe: they took the top 10 search results and answered uncritically based on those. If the results were junk the answer was too. GPT-5 Thinking doesn't do that . Take a look at the thinking trace examples I linked to - in many of them it runs a few searches, evaluates t…

> finds that they're not credible enough to generate an answer

The credibility is one side of the story. In many cases, at least for my curious research, I happen to search for something very niche, so to find at least anything related, an LLM needs to find semantic equivalence between the topic in the query and what the found pages are discussing or explaining.

One recent example: in a flat-style web discussion, it may be interesting to somehow visually mark a reply if the comment is from a user who was already in the discussion (at least GP or GGP). I wanted to find some thoughts or talk about this. I had almost no luck with Perplexity, which probably brute-forced dozens of result pages for semantic equivalence comparison, and I also "was not feeling/getting lucky" with Google using keywords, the AROUND operator, and so on. I'm sure there are a couple of blogs and web-technology forums where this was really discussed, but I'm not sure the current indexing technology is semantically aware at scale.

It's interesting that sometimes Google is still better, for example, when a topic I’m researching has a couple of specific terms one should be aware of to discuss it seriously. Making them mandatory (with quotes) may produce a small result set to scan with my own eyes.

Post reply on HN