Live data from Hacker News

GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

simonwillison.net

261–268 of 268 posts

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#261
post #104

Earlier quoted context omitted.

Yes, this is very much my experience too. Switching to GPT5 Thinking helps a little, but it often misses things that it wouldn't when I was using o3 or o1. As an example, I asked it if there were any incidents involving Botchan in an Onsen. This is a text that is readily available and must have been trained on; in the book, Botchan goes swimming in the onsen, and then is humiliated when the next time he comes back, t…

I feel like the value of my plus subscription went down when they released GPT-5, it feels like a downgrade from o3. But of course OpenAI being not open, there is no way for me to know now.

Likewise.

I'll play devil's advocate and say that I think the Codex-cli included with the plus subscription is pretty good (quality wise). However, after using it, it suddenly told me I couldn't use it for a week out without warning. Claude is a bit more reasonable there.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#262
post #193

Earlier quoted context omitted.

For real. This is what it must have been like living in the early 20th century and hearing people say they prefer a horse to get groceries because it is so much more effort to crank-start a car. I look forward to the age when we gleefully reminisce about the time we had to deal with SEO spam manually.

I look forward to the day AI hype is dead as blockchain.

Most colleagues use AI daily.

It's not going away, ever.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#263
post #249
post #172

Earlier quoted context omitted.

First, you not having to spend the 60 seconds and it means you can parallelize it with something else to get the answer effectively instantly. Second, you're essentially establishing that if an LLM can get it done in less than 60 seconds its better than your manual approach, which is a huge win, as this will get faster!

There's no useful parallelization that could happen during this particular search. This took a couple of iterations of research via ChatGPT, and then reading the results and looking at the referenced sources; the total interaction time with ChatGPT is a similar 60 seconds or so, the main difference is the 3 minutes of waiting for it to generate answers vs. the maybe a couple of seconds for the searches.

What? You can do anything else you want during the search?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#264
post #234
post #190

Earlier quoted context omitted.

Better source validation is one of the main reasons I'm excited about GPT-5 Thinking for this. It would be interesting to try your Gemini prompts against that and see how the results compare.

I've found GPT-5 Thinking to perform worse than o3 did in tasks of a similar nature. It makes more bad assumptions that de-rail the train of thought.

I think the key is prompting, and bound boxing assumptions.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#265

I agree with Simon’s article but I usually think about “research” to mean comparing different kinds of evidence (not just the search part). Like evidence for the effectiveness of Obamacare. Or how some legal case may play out in the courts. Or how much The Critic influenced The Family Guy. Or even what the best way to use X feature of Y library. I’ve found ChatGPT and other LLMS can struggle to evaluate evidence - to…

I tried to use perplexity to find ideal settings for my monitor, it responded with concise list of distinct settings and why. When I investigated the source it was just people guessing and arguing with each other in the Samsung forums, no official or even backed up information. I'd love if it had a confidence rating based on the sources it found or something, but I imagine that would be really difficult to get right.

In the absence of easily found authoritative information from the manufacturer, this would have been my source of information. Internet banter might actually be the best available information.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#266
Even if languages could handle formatting ‘automatically,’ teams would still need style consistency across tools/editors. Formatting isn’t just about syntax, it’s about communication. Maybe the real question is: how much of code style is human preference vs machine necessity?

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#267
post #262
post #193

Earlier quoted context omitted.

I look forward to the day AI hype is dead as blockchain.

Most colleagues use AI daily. It's not going away, ever.

Those colleagues can still use their machine learning autocomplete, I hate the hype, not the tech.

Re: GPT-5 Thinking in ChatGPT (a.k.a. Research Goblin) is good at search

#268
post #61

Earlier quoted context omitted.

As someone who is AI skeptical, there's so many breathless posts like "Jizz-7 Thinking (Good) (Big Balls) can order my morning coffee!" which are a lot of words talking about one person's subjective experience of using some LLM to do one specific thing.

Could you post a selection? It would be intersting to gauge what you mean by breathless. People posting their subjective experience is precisely what a lot of these pieces should be doing, good or bad, their experience is the data they have to contribute.

> People posting their subjective experience is precisely what a lot of these pieces should be doing, good or bad, their experience is the data they have to contribute.

The plural of anecdote is not data. These subjective posts about experiences vibe coding, etc. may be entertaining but if you read 10 of them it doesn't give you an objective view of the state of LLMs. It gives you 10 opinions by 10 people who chose to blog about how they felt using a tool.

Post reply on HN