Live data from Hacker News

AI assisted search-based research works now

simonwillison.net

151–156 of 156 posts

Re: AI assisted search-based research works now

#151

Earlier quoted context omitted.

I think you might as well be saying "robotics fail miserably at the very simple task of jogging around the block, so why should we trust the field to be able to accurately place millions of transistors within a 25cm square of silicon?" His point is that the two tasks are very different at their core, and deep research is better at teasing out an accurate "fuzzy" answer from a swamp of interrelated data, and a data sc…

You seem to misunderstand my previous comment and also the thing being criticized by the post I replied to. I understand that there are fuzzy tasks that AIs/algorithms are terrible at, which seem really simple for a human mind, and this hasn't gone away with the latest generations of LLMs. That's fine and I wouldn't criticize an AI for failing at something like the instructions you describe, for example. However in t…

[deleted]

Re: AI assisted search-based research works now

#152
post #92

Earlier quoted context omitted.

This is just a bad match to the capabilities. What you are actually looking for is analysis, similar in nature to what a data scientist may do. The deep research capabilities are much better suited to more qualitative research / aggregation.

So it's not doing well in things that we can verify/measure, but sure it's doing much better in things we can't measure - except we can't measure them, so we have no idea about how well it is doing actually. The most impressive feature of LLMs stays its ability to impress.

Yup. Like humans.

Re: AI assisted search-based research works now

#153
post #92

Earlier quoted context omitted.

So it's not doing well in things that we can verify/measure, but sure it's doing much better in things we can't measure - except we can't measure them, so we have no idea about how well it is doing actually. The most impressive feature of LLMs stays its ability to impress.

Yup. Like humans.

At least in a liberal society humans matter. Their opinions, judgements, tastes matter. Why should "opinion" (which is not even a real opinion) of a machine matter?

Not to say that we validate whether to trust an opinion of a human expert by them being able to deliver measurably correct judgements, the same thing LLM seem to be not good at.

Re: AI assisted search-based research works now

#154
post #153

Earlier quoted context omitted.

Yup. Like humans.

At least in a liberal society humans matter. Their opinions, judgements, tastes matter. Why should "opinion" (which is not even a real opinion) of a machine matter? Not to say that we validate whether to trust an opinion of a human expert by them being able to deliver measurably correct judgements, the same thing LLM seem to be not good at.

You can't measure the results of your legal advice in most cases. There are all manner of things we can't measure well. We don't throw up our hands and say "then forget it". We do our best with the anecdotes and move forward.

Re: AI assisted search-based research works now

#155
post #89
post #58

Earlier quoted context omitted.

That is an excellent prompt to tuck away in your back pocket and try again future iterations of this technology. It's going to be an interesting milestone when or if any of these systems get good enough at comprehensive research to provide a correct answer.

If you keep the prompt the same at some point the data will appear in training set and we might have answer. So even though today it might be a good check it might not remain as such a good benchmark. I think we need a way to keep updating prompts without increasing complexity in someway to properly verify model improvements. ARC Deep Research anyone?

I had o3 "cheat" yesterday. I tried to demo a Deep Research task to a friend, but o3 managed to find the answer immediately in a Reddit comment I'd made after trying out the same problem previously.

I was still impressed though!

Re: AI assisted search-based research works now

#156

Earlier quoted context omitted.

I'm not a researcher, but don't most researchers these days also upload their work to arXiv? Sure, it's not a journal - but in some fields (Machine Learning, Math) it seems like everyone also uploads their stuff there. So if the models can crawls sites like arXiv, at least there's some decent stuff to be found.

Not outside of ML, physics, and math. Preprints are extremely rare in many (dare I say most) scientific fields, and of course many times you are interested in not the cutting edge work, but the foundational work in a field from the 60s, 70s, or 80s, all of which is locked behind a paywall. Or at least it's supposed to be, and corporate LLMs are not "allowed" to go poking around on sketchy Russian website for non-payw…

I'm assuming many AI companies are probably scraping all the PDFs from the "shadow libraries" of the world that have done some of the work of liberating these papers from behind their paywalls. Obviously it's legally unsettled territory right now...
Post reply on HN