Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

231–240 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#231

I just can't get over how gleeful the author sounds in wasting compute on "an often unreasonable amount of work to search the internet and figure out an answer." Is that the goal? Send this thing off on a wild goose chase, and hope it comes back with the right answer no matter the cost?

Uh, people have wasted entire lifetimes chasing wild goose. Newton and Einstein both spent the latter halves of their lives :( despite being geniuses.

I think the primary difference would be they didn't waste billions of dollars in their research.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#233

I was curious how much revenue a podcast I listen to makes. The podcast was started by two local comedians from Phoenix, AZ. They had no following when they started and were both in their late 30's. The odds were stacked against them, but they rank pretty high now on the the Apple charts now. I looked into years ago and couldn't find a satisfying answer, but GPT-5 went off, did an "unreasonable" amount of research, c…

What was the range?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#234
post #190

Earlier quoted context omitted.

I asked gemini to do a deep research on the role of healthcare insurance companies in the decline of general practicioners in the Netherlands. It based its premise mostly on blogs and whitepapers on company websites, who's job it is to sell automation-software. AI really needs better source-validation. Not just to combat the hallucination of sources (which gemini seems to do 80% of the time), but also to combat low q…

Better source validation is one of the main reasons I'm excited about GPT-5 Thinking for this. It would be interesting to try your Gemini prompts against that and see how the results compare.

I've found GPT-5 Thinking to perform worse than o3 did in tasks of a similar nature. It makes more bad assumptions that de-rail the train of thought.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#235
post #222
post #77

Earlier quoted context omitted.

I was amused that it used the neologism 'steel-man' -- redundantly, too.

I'm a bit confused, how is it redundant here? It's trying to make the best possible argument from one side that seems to be wrong. Instead of taking the argument at face value, it takes the most charitable understanding of it (not requiring that it happened before, but some parts where perhaps inspired during later revisions) and tries to argue that case.

'strongest “steel-man” case' is the same thing as strongest case; “steel-man” adds nothing.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#236

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

I found your article interesting and it is relevant to the discussion. To be honest, while I think GPT could have performed better here, I think there is something to be said about this: There is value in pruning the search tree because the deeper nodes are usually not reputable. I know you have cause to believe that "Wilhelm Matzat" is reputable but I don't think it can be assumed generally. If you were to force GPT…

My contention is if it’s going to just give me a Wikipedia summary, I can do that myself. I just have greater expectations of “PhD” level intelligence.

If we’re going to claim to it is PhD level it should be able to do “deep” research AND think critically about source credibility, just as a PhD would. If it can’t do that they shouldn’t brand it that way.

Also it’s not like I’m taking Matzat’s word for anything. I can read the primary source documents myself! He’s also hardly an obscure source, he’s just not listed on Wikipedia.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#237
post #222
post #77

Earlier quoted context omitted.

I was amused that it used the neologism 'steel-man' -- redundantly, too.

I'm a bit confused, how is it redundant here? It's trying to make the best possible argument from one side that seems to be wrong. Instead of taking the argument at face value, it takes the most charitable understanding of it (not requiring that it happened before, but some parts where perhaps inspired during later revisions) and tries to argue that case.

The question I asked was intentionally trying to see if GPT would go for it or just give me the answer it thought I wanted, but it did a pretty decent job at not just saying “you’re absolutely right” etc. Myself I don’t believe there to be much influence between the two.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#238

Earlier quoted context omitted.

I found your article interesting and it is relevant to the discussion. To be honest, while I think GPT could have performed better here, I think there is something to be said about this: There is value in pruning the search tree because the deeper nodes are usually not reputable. I know you have cause to believe that "Wilhelm Matzat" is reputable but I don't think it can be assumed generally. If you were to force GPT…

My contention is if it’s going to just give me a Wikipedia summary, I can do that myself. I just have greater expectations of “PhD” level intelligence. If we’re going to claim to it is PhD level it should be able to do “deep” research AND think critically about source credibility, just as a PhD would. If it can’t do that they shouldn’t brand it that way. Also it’s not like I’m taking Matzat’s word for anything. I can…

I suggest ignoring the "PhD level intelligence" marketing hype.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#239
post #233

I was curious how much revenue a podcast I listen to makes. The podcast was started by two local comedians from Phoenix, AZ. They had no following when they started and were both in their late 30's. The odds were stacked against them, but they rank pretty high now on the the Apple charts now. I looked into years ago and couldn't find a satisfying answer, but GPT-5 went off, did an "unreasonable" amount of research, c…

What was the range?

$2.3M–$3.7M per year gross revenue:

Putting it together (podcast‑only)

Ads (STM only): 0.85 – $0.95M

Memberships: $1.5–$2.6M/yr

Working estimate (gross, podcast‑only): $2.3M – $3.7M per year for ads + its share of memberships. Mid‑case lands near $2.9M/yr gross; after typical platform/processing fees and less‑than‑perfect ad sell‑through, a net in the low‑to‑mid $2Ms seems plausible.

---

Quick answers

How many listeners: 175K downloads per episode, ~1.4–2.1M monthly downloads, demographics skew U.S., median age ~36.

How much revenue? $2.3M– $3.7M/yr gross from ads + memberships attributable to the STM show

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#240

I just can't get over how gleeful the author sounds in wasting compute on "an often unreasonable amount of work to search the internet and figure out an answer." Is that the goal? Send this thing off on a wild goose chase, and hope it comes back with the right answer no matter the cost?

Playing with tools like this is how I learn to use them.
Post reply on HN