Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

141–150 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#141

Earlier quoted context omitted.

I asked why it's popping up over and over again in new today. I wouldn't have commented otherwise.

From https://en.wikipedia.org/wiki/Hacker_News : "Hacker News (HN) is a social news website" From https://en.wikipedia.org/wiki/Social_news_website : "A social news website is a website that features user-posted stories. Such stories are ranked based on popularity, as voted on by other users of the site or by website administrators." The article was recently published, users on HN submitted the article. Other users t…

It's 8 hours from time of post no? So timezones don't really affect anything here or am I missing something?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#142

Earlier quoted context omitted.

Simon's writing is consistently either highly practical, or extremely high quality, or both. What's your reference frame to call it "bad" - your own comments?

Spending thousands of words to essentially say "ChatGPT's search feature works pretty well now" with mundane examples like finding UK cake pop availability or identifying buildings from train windows. This has been done before by less capable models - it's just a rehash. Should we expect newer models getting worse? The breathless "Research Goblin" framing and detailed play-by-play of basic web searches feels like pad…

Yes this feels very AI like too. A ton of prose for very little substance lol.

I skipped half the article to get to the point, went back and re-read and didn't miss much.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#143
post #79

I do miss the earlier "heavy" models that had encyclopedic knowledge vs the new "lighter" models that rely on web search. Relying on web search surfaces a shallow layer of knowledge (thanks to SEO and all the other challenges of ranking web results) vs having ingested / memorized basically the entirety of human written knowledge beyond what's typically reachable within the first 10 results of a web search (eg: digiti…

I think this is partially something I have felt myself as well. It would be interesting if these lighter web search models would highlight the distinction between information that has been seen elsehwere vs information that is novel for each page? Like, a view that lets me look at the things that have been asserted and see how many of the different pages show those facts asserted (vs unmentioned vs contradicted).

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#145
post #83

Yes, "GPT-5 with thinking" is great at search, but it's horrible that it shows "Network connection lost. Attempting to reconnect..." after you switch away from the app for even just a few seconds before coming back. It's going to take a minute, so why do I need to keep looking at it and can't go read some more Wikipedia in the mean time? This is insanely user hostile. Is it just me who encounters this? I'm on Plus pl…

Ive always had this happen too, super annoying. Android.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#146

Earlier quoted context omitted.

From your blog you appear to be a Georgist or inspired by Georgist socialism. And given that you appear to have a business and blog related to these subjects, you give the impression that you're a sort of activist for Georgism. I.e. not just researching it by trying to advance it. So just zooming out, that's not the right sort of setup for being an impartial researcher. And in your blog post your disagreements come o…

You misunderstand. I’m indeed a Georgist, and I discovered that a popular Georgist narrative was exaggerated! The findings of the historically verifiable primary source documents contradicted a prevailing narrative based on the Silagi paper. The Silagi paper is pro Georgist! But it’s exaggerated! The literature — the primary source documents — do not in fact support a maximalist Georgist case! This is what I have bee…

To clarify I am not accusing you of that. I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous. For example whether a given policy is Georgist.

But often people who believe in a given doctrine will see differences as more important than they objectively are. For example, just to continue with socialism, it's common for socialist believers to argue that this or that country is or isn't socialist in a way that disagrees with mainstream historians.

I'm sure there are other examples, for example people disagreeing about which bands are punk or hardcore. A music historian would likely cast a wider net. Fans who don't listen to many other types of music might cast a very narrow net.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#147

[flagged]

HN is very cult-of-personality based. People see SimonW they upvote without reading, while at the same time a much better article could be posted on the same topic and get zero traction. Not trying to single Simon out here, I generally find his posts good, just a statement of the herdthink and cognitive laziness of this community (and humans in general, to be fair).

I mean if you have a better idea for how to assign your attention, then I am all ears. :]

I'd say trust is a pretty reasonable way to assign attention.

I guess the fairest way might theoretically be to require everything to be submitted anonymously, with maybe authorship (maybe submissionship) only being revealed after some assigned period?

This is better for the incubants, but would require a huge amount of energy compared to "Oh, simon finds this interesting, I'll take a looksy".

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#148

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

How are we feeling about the usage of the word research to indicate feature sets in LLMs? Is it truly representative of research? How does it compare to the colloquial “do your research” refrain used often during US election years?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#149

Earlier quoted context omitted.

You misunderstand. I’m indeed a Georgist, and I discovered that a popular Georgist narrative was exaggerated! The findings of the historically verifiable primary source documents contradicted a prevailing narrative based on the Silagi paper. The Silagi paper is pro Georgist! But it’s exaggerated! The literature — the primary source documents — do not in fact support a maximalist Georgist case! This is what I have bee…

To clarify I am not accusing you of that. I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous. For example whether a given policy is Georgist. But often people who believe in a given doctrine will see differences as more important than they objectively are. For example, just to continue with socialism, it's common for socialist beli…

Okay, so let me break it down for you:

The Silagi paper makes a factual claim. The Silagi paper claims that there was only one significant tax in the German colony of Kiatschou, a single tax on land.

The direct primary sources reveal that this is not the case. There were multiple taxes, most significantly large tariffs. Additionally there were two taxes on land, not one -- a conventional land value tax, and a "land increment" or capital gains tax.

These are not minor distinctions. These are not matters of subjective opinions. These are clear, verifiable, questions of fact. The Silagi paper does not acknowledge them.

ChatGPT, in the early trials I graded, does not even acknowledge the German primary sources. You keep saying that I am upset it doesn't agree with me.

I am saying the chief issue is that ChatGPT does not even discover the relevant primary sources. That is far more important than whether it agrees with me.

> For example, just to continue with socialism, it's common for socialist believers to argue that this or that country is or isn't socialist in a way that disagrees with mainstream historians.

Notice you said "historians." Plural. I expect a proper researcher to cite more than ONE paper, especially if the other papers disagree, and even if it has a preferred narrative, to at least surface to me that there is in fact disagreement in the literature, rather than to just summarize one finding.

Also, if the claims are being made about a piece of German history, I expect it to cite at least one source in German, rather than to rely entirely on one single English-language source.

The chief issue is that ChatGPT over-cites one single paper and does not discover primary source documents. That is the issue. That is the only issue.

> I am saying you are seeing distinctions as more important than the rest of the literature and concluding that the literature is erroneous.

And I am saying that ChatGPT did not in fact read the "rest of the literature." It is literally citing ONE article, and other pieces that merely summarize that same article, rather than all of the primary source documents. It is not in fact giving me anything like an accurate summary of the literature.

I am not saying "The literature is wrong because it disagrees with me." I am saying "one paper, the only one ChatGPT meaningfully cites, is directly contradicted by the REST of the literature, which ChatGPT does not cite."

A truly "research grade" or "PhD grade" intelligence would at the very least be able to discover that.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#150

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

How are we feeling about the usage of the word research to indicate feature sets in LLMs? Is it truly representative of research? How does it compare to the colloquial “do your research” refrain used often during US election years?

Well I will just need to start saying “critical thinking”? Or some other term?

I have a liberal arts background. So I use the term research to mean gathering evidence, evaluating its trustworthiness and biases, and avoiding related thinking errors related to evaluating evidence (https://thedecisionlab.com/biases).

LLMs can fall prey to these problems as well. Usually it’s not just “reasoning” that gives you trouble. It’s the reasoning about evidence. I see this with Claude Code a lot. It can sometimes create some weird code, hallucinating functionality that doesn’t exist, all because it found a random forum post.

I realize though that the term is pretty overloaded :)

Post reply on HN