Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

211–220 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#211

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

This isn't really how you described. You have an opinion that conflicts with the research literature. You published a blog about that opinion, and you want ChatGPT to say you're to accept your view. Your view is grinding a political axe and I don't think you're in a position to objectively assess whether ChatGPT failed in this case.

Hmm, I suspect if ChatGPT would pay more attention to the German sources, they would perhaps find that supposedly right answer?

I wonder if asking ChatGPT in German would make a difference.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#212

Earlier quoted context omitted.

Yea this isn't really a chat gpt problem as a source credibility problem no?

It’s mostly that it was not citing verifiable - and available online - primary source documents, the way I would expect an actual researcher investigating this question would. This is relevant when it is billed as "Research Grade" or "PhD" level intelligence. I expect a PhD level researcher to find the German-language primary sources.

Especially since ChatGPT speaks fluent German.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#213

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

More recently, I find ChatGPT to become increasingly unreliable. It makes up almost every second answer, forgets context, or is just downright wrong. Maybe I am used these days more and more to dump huge texts for context into the prompt, as aistudio allows me. Maybe ChatGPT isn't as good as with such information. Gemini/Aistudio will stay on track even with 300k tokens consumed, it just needs a little nudge here and there.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#214
Has Deep Research been removed? I have a Pro subscription and just today noticed Deep Research is no longer shown as an option. In any case I’ve found using GPT-5 Thinking, and especially GPT-5 Pro with web search more useful than DR used to be.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#215

Earlier quoted context omitted.

I wonder how all this will really change the web. In your manual mode, you a human, are viewing and visiting webpages, but if one never needs to and always interacts with the web through an agent, what does the web need to look like, and will people even bother making websites? Interesting times ahead.

I’ve been thinking about this as well. Instead of making websites, maybe people will make something else, like some future version of MCP tools/servers? E.g. a restaurant could have an “MCP tool” for checking opening hours, reserving a table, etc.

Same. Websites won't disappear but may become niche or something of the past. Why create a new UI for your new service when you can plug into a "universal" personal agent AI.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#216

Earlier quoted context omitted.

To clarify I am not accusing you of that. I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous. For example whether a given policy is Georgist. But often people who believe in a given doctrine will see differences as more important than they objectively are. For example, just to continue with socialism, it's common for socialist beli…

Okay, so let me break it down for you: The Silagi paper makes a factual claim. The Silagi paper claims that there was only one significant tax in the German colony of Kiatschou, a single tax on land. The direct primary sources reveal that this is not the case. There were multiple taxes, most significantly large tariffs. Additionally there were two taxes on land, not one -- a conventional land value tax, and a "land i…

[deleted]

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#217
I was curious how much revenue a podcast I listen to makes. The podcast was started by two local comedians from Phoenix, AZ. They had no following when they started and were both in their late 30's. The odds were stacked against them, but they rank pretty high now on the the Apple charts now.

I looked into years ago and couldn't find a satisfying answer, but GPT-5 went off, did an "unreasonable" amount of research, cross referenced sources and provided an incredibly detailed answer and a believable range.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#218

I was curious how much revenue a podcast I listen to makes. The podcast was started by two local comedians from Phoenix, AZ. They had no following when they started and were both in their late 30's. The odds were stacked against them, but they rank pretty high now on the the Apple charts now. I looked into years ago and couldn't find a satisfying answer, but GPT-5 went off, did an "unreasonable" amount of research, c…

> an incredibly detailed answer and a believable range.

Recently, it started returning even more verbose answers. The absolute bullshit research paper that Google Gemini gives you is what turned me away from using it there. Now, chatGPT also seems to go for more verbose filler rather than actually information. It is not as bad as Gemini, but I did notice.

It makes me wonder if people think the results are more credible with verbose reports like that. Even if it actually obfuscates the information you asked it to track down to begin with.

I do like how you worded it as a believable range, rather than an accurate one. One of the things that makes me hesitant to use deep research for anything but low impact non-critical stuff is exactly that. Some answers are easier to verify than others, but the way the sources are presented, it isn't always easy to verify the answers.

Another aspect is that my own skills in researching things are pretty good, if I may so myself. I don't want them to atrophy, which easily happens with lazy use of LLMs where they do all the hard work.

A final consideration came from me experimenting with MCPs and my own attempts at creating a deep research I could use with any model. No matter the approach I tried, it is extremely heavy on resources and will burn through tokens like no other.

Economically, it just doesn't make sense for me to run against APIs. Which in my mind means it is heavily subsidized by openAI as a sort of loss-leader. Something I don't want to depend on just to find myself facing a price hike in the future.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#219

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

But the really tricky thing is, that sometimes it _is_ these kinds of forums where you find the best stuff.

When LLMs really started to show themselves, there was a big debate about what is truth, with even HN joining in on heated debates on the number of sexes or genders a dog may have and if it was okay or not for ChatGPT to respond with a binary answer.

On one hand, I did found those discussions insufferable, but the deeper question - what is truth and how do we automated the extraction of truth from corpora - is super important and somehow completely disappeared from the LLM discourse.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#220

Earlier quoted context omitted.

To clarify I am not accusing you of that. I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous. For example whether a given policy is Georgist. But often people who believe in a given doctrine will see differences as more important than they objectively are. For example, just to continue with socialism, it's common for socialist beli…

Okay, so let me break it down for you: The Silagi paper makes a factual claim. The Silagi paper claims that there was only one significant tax in the German colony of Kiatschou, a single tax on land. The direct primary sources reveal that this is not the case. There were multiple taxes, most significantly large tariffs. Additionally there were two taxes on land, not one -- a conventional land value tax, and a "land i…

I think we’re talking past each other a bit. My concern is that your personal assessment of whether the tax is "significant" is being treated as settled fact. That’s the same kind of issue I flagged earlier. Reasonable people can disagree here without that disagreement implying a "pathological failure."

I hear you that this is about finding sources, but even perfect coverage of primary sources wouldn’t remove the need for judgment. We’d still have to define what counts as "Georgian," "inspired by George," and "significant" as a tax. Those are contestable choices. What you have is a thesis about the evidence—potentially a strong one—but it isn’t an indisputable fact.

On sourcing: I’m aware ChatGPT won’t surface every primary source, and I’m not sure that should be the default goal. In many fields (e.g., cancer research), the right starting point is literature reviews and meta-analyses, not raw studies. History may differ, but many primary sources live offline in archives, and the digitized subset may not be representative. Over-weighting primary materials in that context can mislead. Primary sources also demand more expertise to interpret than secondary syntheses—Wikipedia itself cautions about this: https://en.wikipedia.org/wikiWikipedia:Identifying_and_using...

To be clear, I’m not saying you’re wrong about the tax or that Silagi is right. I’m saying that framing this as a “pathological failure” overstates the situation. What I see is a legitimate disagreement among competent researchers.

Post reply on HN