Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

251–260 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#251

Earlier quoted context omitted.

For real. This is what it must have been like living in the early 20th century and hearing people say they prefer a horse to get groceries because it is so much more effort to crank-start a car. I look forward to the age when we gleefully reminisce about the time we had to deal with SEO spam manually.

The thing is, the idea of cars cars being a plus for humanity is still debatable to this day. We won a lot in some areas but lost an awful lot in others.

That's a ridiculously cynical claim. And even if it were true, it still misses the fact that horses and the entire economy built around them stood no chance in the end. Today we are again in the sitation where many people who are alive and working today will have to consider getting reeducated if they don't want to be left behind, whether they want it or not.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#252
post #243
post #166

Earlier quoted context omitted.

I suggest trying that experiment again but picking the hardest of my examples to answer with Google, not the easiest.

Not sure which is the hardest, but sure, let's try them all. * Bouncy people mover. Some Google searching turns up the SFO article that you liked. Trying to pin down the exact dates is harder. ChatGPT maybe did narrow down the time frame quicker than I could through a series of Google searches, * The picture of the building. Go to Google lens, paste in the image, less than a second later I get results. Of course, the…

This is a fair analysis, thanks for taking the time.

As far as I can tell the Google + Wikipedia solution gets the name of Cambridge University wrong: Wikipedia lists it as "The Chancellor, Masters and Scholars of the University of Cambridge" whereas GPT-5 correctly verified it to be "The Chancellor, Masters, and Scholars of the University of Cambridge" (note that extra comma) as listed on https://www.cam.ac.uk/about-the-university/how-the-universit...

I tried to reverse engineer the system prompt in the cake pop conversation https://chatgpt.com/share/68bc71b4-68f4-8006-b462-cf32f61e7e... purely because I got annoyed at it for answering "haha I believe you" - I particularly disliked the lower case "haha" because I've seen it switch to lower case (even the lower case word "i") in the past and I wanted to know what was causing it to start talking in the same way that Sam Altman tweets.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#253

I find ChatGPT to be great at research too-but there are pathological failure modes where it is biased to shallow answers that are subtly wrong, even when definitive primary sources are readily available online: https://www.fortressofdoors.com/researchers-beware-of-chatgp...

This doesn't tell us much. I don't know why you would expect ChatGPT to do original PhD research. It's a general product that will trust already published research. That doesn't meat that GPT-5 can't do PhD research, when given the right sources.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#255
post #240

I just can't get over how gleeful the author sounds in wasting compute on "an often unreasonable amount of work to search the internet and figure out an answer." Is that the goal? Send this thing off on a wild goose chase, and hope it comes back with the right answer no matter the cost?

Playing with tools like this is how I learn to use them.

I can definitely understand that. I just have a hard time justifying the value of learning this way at scale. Just the energy costs alone are out of hand.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#256

Earlier quoted context omitted.

Uh, people have wasted entire lifetimes chasing wild goose. Newton and Einstein both spent the latter halves of their lives :( despite being geniuses.

I think the primary difference would be they didn't waste billions of dollars in their research.

Isaac Newton dedicated over thirty years to the study and practice of alchemy, writing over one million words on the subject. Comparable in scale to his writings on mathematics and physics.

I'd rather GDP be $1B smaller right now if it meant that Newton had spent another 30 years on physics and math.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#257

Earlier quoted context omitted.

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

I asked gemini to do a deep research on the role of healthcare insurance companies in the decline of general practicioners in the Netherlands. It based its premise mostly on blogs and whitepapers on company websites, who's job it is to sell automation-software. AI really needs better source-validation. Not just to combat the hallucination of sources (which gemini seems to do 80% of the time), but also to combat low q…

When using AI models through Kagi Assistant you can tweak the searches the LLM does with your Kagi settings (search only academic, block bullshit websites and such) which is nice. And I can chose models from many providers.

No API access though so you're stuck talking with it through the webapp.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#258
post #252
post #243

Earlier quoted context omitted.

Not sure which is the hardest, but sure, let's try them all. * Bouncy people mover. Some Google searching turns up the SFO article that you liked. Trying to pin down the exact dates is harder. ChatGPT maybe did narrow down the time frame quicker than I could through a series of Google searches, * The picture of the building. Go to Google lens, paste in the image, less than a second later I get results. Of course, the…

This is a fair analysis, thanks for taking the time. As far as I can tell the Google + Wikipedia solution gets the name of Cambridge University wrong: Wikipedia lists it as "The Chancellor, Masters and Scholars of the University of Cambridge" whereas GPT-5 correctly verified it to be "The Chancellor, Masters, and Scholars of the University of Cambridge" (note that extra comma) as listed on https://www.cam.ac.uk/about…

But that's an Oxford comma! You can't use that when describing Cambridge.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#259

Earlier quoted context omitted.

I’ve been thinking about this as well. Instead of making websites, maybe people will make something else, like some future version of MCP tools/servers? E.g. a restaurant could have an “MCP tool” for checking opening hours, reserving a table, etc.

I hope none of this happens and web stays readable and indexable.

I sure hope it stays readable, but it seems like it would only become more indexable with machine friendly formats.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#260

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

Seems like the right outcome was had, by reviewing sources. I wish it went one step further and loaded those source pages and scroll/highlight the snippets where it pulled information from. That way we can easily double check at least some aspects of it's response, and content+ads can be attributed to the publisher.
Post reply on HN