Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

201–210 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#201
post #157

I guess the part where I'm still skeptical are: Google is also still pretty good at search (especially if I avoid the AI summary with udm=14). I'll take one of your examples: Britannica to seed Wikipedia. I searched for "wikipedia encyclopedia brtannica". In less than 1 second, I got search results back. I spend maybe 30 seconds scanning the page; past the Wikipedia article on Encyclopedia Britannica, past the Encycl…

Agree. I tried the first 3 examples:

* "Rubber bouncy at Heathrow removal" on Google had 3 links, including the one about SFO from which chatGPT took a tangent. While ChatGPT provided evidence for the latest removal date being of 2024, none was provided for the lower bound. I saw no date online either. Was this a hallucination?

* A reverse image lookup of the building gave me the blog entry, but also an Alamy picture of the Blade (admittedly this result can have been biased by the fact the author already identified the building as the blade)

* The starbucks pop Google search led me to https://starbuckmenu.uk/starbucks-cake-pop-prices/. I will add that the author bitching to ChatGPT about ChatGPT hidden prompts in the transcript is hilarious.

I get why people prefer ChatGPT. It will do all the boring work of curating the internet for you, to privde you with a single answer. It will also hallucinate every now and then but that seems to be a price people are willing to pay and ignore, just like the added cost compared to a single Google search. Now I am not sure how this will evolve.

Back in the days, people would tell you to be weary of the Internet and that Wikipedia thing, and that you could get all the info you need from a much more reliable source at the library anyways, for a fraction of the cost. I guess that if LLMs continue to evolve, we will face the same paradigm shift.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#202

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

> I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to understand the biases behind sources - ie taking data from a sketchy think tank as gospel. This is what I keep finding, it mostly repeats surface level "common knowledge." It usually take a few back and forths to get to whether or not something is actually true - asking for the numbers, asking for the sources, asking for the excerpt from the s…

> It usually take a few back and forths to get to whether or not something is actually true

This cuts both ways. I have yet to find an opinion or fact I could not make chatgpt agree with as if objectivly true. Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. (Which some already are in 1 way or another)

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#203
I agree, I've found it very useful for search tasks that involve putting a few pieces together. For example, I was looking for the floor plan of an apartment building I was considering moving to in a country I'm not that familiar with. Google: Found floor plans on architect's sites - but only ones of rejected proposals.

ChatGPT: Gave me the planning proposal number of the final design, with instructions to the government website where I can plug that in to get floor plans and other docs.

In this case, ChatGPT was so much better at giving me a definitive source than Google - instead of the other way around.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#205

Earlier quoted context omitted.

> I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to understand the biases behind sources - ie taking data from a sketchy think tank as gospel. This is what I keep finding, it mostly repeats surface level "common knowledge." It usually take a few back and forths to get to whether or not something is actually true - asking for the numbers, asking for the sources, asking for the excerpt from the s…

> It usually take a few back and forths to get to whether or not something is actually true This cuts both ways. I have yet to find an opinion or fact I could not make chatgpt agree with as if objectivly true. Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. (Which some already are in 1 way or another)

I'm not sure teaching it in school is actually going to help. Most people will tell you that of course you need to look at primary sources to verify claims - and then turn around and believe the first thing they here from LLM, Redditor, Wiki article, etc. Even worse, many people get openly hostile to the idea that people should verify claims - "what, you don't believe me?"/"everyone here has been telling you this is true, do you have any evidence it isn't?"/"oh, so you think you know better?"

There was a recent discussion about Wikipedia here recently where a lot of people who are active on the site argued against people taking the claims there with a grain of salt and verifying the accuracy for themselves.

We can teach these things until the cows come home, but it's not going to make a difference if people say it's a good idea and then immediately do the opposite.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#206

Earlier quoted context omitted.

I wonder how all this will really change the web. In your manual mode, you a human, are viewing and visiting webpages, but if one never needs to and always interacts with the web through an agent, what does the web need to look like, and will people even bother making websites? Interesting times ahead.

I’ve been thinking about this as well. Instead of making websites, maybe people will make something else, like some future version of MCP tools/servers? E.g. a restaurant could have an “MCP tool” for checking opening hours, reserving a table, etc.

I hope none of this happens and web stays readable and indexable.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#207

Slightly off topic but chatGPT’s refusal to visually identify people, including dead historical personalities, has been a big let down for me. I can paste in an image of JFK and it will refuse to tell me who it is.

I think it makes sense? Given the vast "knowledge" of ChatGPT it'd be a perfect doxxing tool with the deep research. To straight-up refuse any identification is I think a better idea than to try to circumvent it with arbitrary limitations? However, having tried it now myself. Uploading the profile picture of Gauchy and asking it who this person is in the image made it refuse, even after asking who it is. But starting…

I understand the legal motivation behind a blanket ban, but what's the point of having artificial "intelligence" if the model can't contextualize the request? Any intelligent model would be able to figure out that JFK is not under any threat of being doxxed

I legitimately had to ask Reddit for answers because I saw a picture of historical figures where I recognized 3 of the 4 people, but not the 4th. That 4th person has been dead for 78 years. Google Lens, and ChatGPT both refused to identify the person - one of the leading scientists of the 20th century.

You can't really build a tool that you claim can be used as a learning tool but can't identify people without contextualizing the request.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#209

Earlier quoted context omitted.

> I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to understand the biases behind sources - ie taking data from a sketchy think tank as gospel. This is what I keep finding, it mostly repeats surface level "common knowledge." It usually take a few back and forths to get to whether or not something is actually true - asking for the numbers, asking for the sources, asking for the excerpt from the s…

> It usually take a few back and forths to get to whether or not something is actually true This cuts both ways. I have yet to find an opinion or fact I could not make chatgpt agree with as if objectivly true. Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap. (Which some already are in 1 way or another)

> Knowing how to trigger (im)partial thought is a skill in and of itself and something we need to be teaching in school asap.

You are very optimistic.

Look at all other skills we are trying to teach in school. 'Critical thinking' has been at the top of nearly every curriculum you can point a finger at for quite a while now. To minimal effect.

Or just look at how much math we are trying to teach the kids, and what they actually retain.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#210

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

FWIW GPT-5 (and o3, etc.) is one of the most critical-minded LLMs out there. If you ask for information which is e.g. academic or technical it would cite information and compare different results, etc, without any extra prompt or reminder. Grok 4 (at the initial release) was just reporting information in the articles it found without any analysis. Claude Opus 4 also seems bad: I asked it to give a list of JS librarie…

> So GPT-5 is really good in comparison. Maybe not perfect in all situations, but perhaps better than an average human

Alas, the average human is pretty bad at these things.

Post reply on HN