Live data from Hacker News

AI assisted search-based research works now

simonwillison.net

91–100 of 156 posts

Re: AI assisted search-based research works now

#91

The various deep research products don't work well for me. For example I asked these tools yesterday, "How many unique NFL players were on the roster for at least one regular season game during the 2024 season? I'd like the specific number not a general estimate." I as a human know how to find this information. The game day rosters for many NFL teams are available on many sites. It would be tedious but possible for m…

I used Google AI Studio instead of Google Gemini App because it provides references to the search results.

Google AI Studio gave me an exact answer of 2227 as a possible answer and linked to these comments because there is a comment further down which claims that is the exact answer. The comment was 2 hours old when I did the prompt.

It also provided a code example of how to find it using the python nfl data library mentioned in one of the comments here.

Re: AI assisted search-based research works now

#92

The various deep research products don't work well for me. For example I asked these tools yesterday, "How many unique NFL players were on the roster for at least one regular season game during the 2024 season? I'd like the specific number not a general estimate." I as a human know how to find this information. The game day rosters for many NFL teams are available on many sites. It would be tedious but possible for m…

This is just a bad match to the capabilities. What you are actually looking for is analysis, similar in nature to what a data scientist may do. The deep research capabilities are much better suited to more qualitative research / aggregation.

So it's not doing well in things that we can verify/measure, but sure it's doing much better in things we can't measure - except we can't measure them, so we have no idea about how well it is doing actually. The most impressive feature of LLMs stays its ability to impress.

Re: AI assisted search-based research works now

#93
post #20

The main "real-world" use cases for AI use for now have been: - shooting buildings in Gaza https://apnews.com/article/israel-palestinians-ai-weapons-43... - compiling a list of information on Government workers in US https://www.msn.com/en-us/news/politics/elon-musk-s-doge-usi... - creating a few losy music videos I'd argue we'd be better off SLOWING DOWN with that shit

You seem ideology motivated instead of truth motivated which makes you untrustworthy.

So give some citations of other notable uses?

Re: AI assisted search-based research works now

#94
The most impressive demos of these tools always involve technical tasks where the user already knows enough to verify accuracy. But for the average person asking about health issues, legal questions, or historical facts? It's basically fancy snake oil - confident-sounding BS that people can't verify. The real breakthrough would be systems that are actually trustworthy without human verification, not slightly better BS generators. True AI research breakthroughs would admit uncertainty and provide citations for everything, not fake certainty like these tools do.

Re: AI assisted search-based research works now

#95
post #39
post #20

The main "real-world" use cases for AI use for now have been: - shooting buildings in Gaza https://apnews.com/article/israel-palestinians-ai-weapons-43... - compiling a list of information on Government workers in US https://www.msn.com/en-us/news/politics/elon-musk-s-doge-usi... - creating a few losy music videos I'd argue we'd be better off SLOWING DOWN with that shit

Programming is not real world?

I said "the main use cases", not "the little toys to distract and amuse engineers while the ruin the environment with CO2 emissions"

Re: AI assisted search-based research works now

#96
This is surprising. o3 produces incredible amount of hallucinations for me, and there are lots of reddit threads about it. I've had to roll back to another model because it just swamps everything in made up facts. But sometimes it is frighteningly smart. Reading its output sometimes feels like I'm missing IQ points.

Re: AI assisted search-based research works now

#98

The various deep research products don't work well for me. For example I asked these tools yesterday, "How many unique NFL players were on the roster for at least one regular season game during the 2024 season? I'd like the specific number not a general estimate." I as a human know how to find this information. The game day rosters for many NFL teams are available on many sites. It would be tedious but possible for m…

This is just a bad match to the capabilities. What you are actually looking for is analysis, similar in nature to what a data scientist may do. The deep research capabilities are much better suited to more qualitative research / aggregation.

Your logic is.... strange...

Because it failed miserably at a very simple task of looking through some scattered charts, the human asking should blame themselves for this basic failure and trust it to do better with much harder and more specialized tasks?

Re: AI assisted search-based research works now

#99
A common google searching thing I counter have is something like this:

I need to get from A to B via C via public transport in a big metropolis.

Now C could be one of say 5 different locations of a bank branch, electronics retailer, blood test lab or whatever, so there's multiple ways of going about this.

I would like a chatbot solution that compares all the different options and lays them out ranked by time from A to B. Is this doable today?

Re: AI assisted search-based research works now

#100
post #89
post #58

Earlier quoted context omitted.

That is an excellent prompt to tuck away in your back pocket and try again future iterations of this technology. It's going to be an interesting milestone when or if any of these systems get good enough at comprehensive research to provide a correct answer.

If you keep the prompt the same at some point the data will appear in training set and we might have answer. So even though today it might be a good check it might not remain as such a good benchmark. I think we need a way to keep updating prompts without increasing complexity in someway to properly verify model improvements. ARC Deep Research anyone?

Wouldn't somebody need to answer the question below? Or do you mean the discussion of its weakness might somehow make it stronger the next time it's trained?
Post reply on HN