Live data from Hacker News

The AI Scientist: Towards Automated Open-Ended Scientific Discovery

sakana.ai

21–30 of 144 posts

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#22

To produce scientific work, one needs certain raw materials: 1. Data 2. Access to past works Once you have these, only then can discoveries can be made, and papers be written. How does this software get these? I am assuming they have to be provided up-front to the software for each job.

You need a meaningful cost function

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#23

To produce scientific work, one needs certain raw materials: 1. Data 2. Access to past works Once you have these, only then can discoveries can be made, and papers be written. How does this software get these? I am assuming they have to be provided up-front to the software for each job.

It will analyse what it is given, but it will not have the ability to say, "hang on, these results are interesting, I wonder what will happed if I pour a different liquid into the drum and spin it at the same speed?" LLMs, especially the latest ones, are decent at analysis of input, but disappointing at producing creative output. I am running a series of experiments using Gemini 1.5 and found it capable of producing good results if you stay away from "write me an academic paper on subject X" or "write me a novel". On the other hand, if you ask it to summarise text, extract particular information, it is fast and arguably good, but not necessarily great. It will miss things and miscategorise them requiring a human being to check its output. At this point, you may just as well do the job yourself. LLMs are still not very good, despite what their fans are saying. They are clever, as in a "clever trick" not as in a "clever human being".

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#24
post #9

Clarkesworld sci-fi magazine temporarily closed submissions due to low quality AI spam. I'm sure the irony will not be lost on them if ML journals are the next victims.

AI is a boon to Ph.D. factories offering fasttrack to a degree.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#25

But OpenAI said LLMs can't innovate until human-level reasoning and long-term agenthood is solved. [1] Referring to their precious 5 stages to classify AI before it reaches the scary "beyond" levels of intelligence... presumably at that point they get the feds involved to reg cap the field, so genuine is the fear of the pace they've set. It's clear OpenAI is a hype company knocking over glass bottle stacks at its own…

It would also be sad to see the scientific system destroyed by a wave of automatically generated papers that no human has the capacity to verify.

It's not hard to generate ideas, it's hard to generate reliable and relevant ideas. Such AI science generators are destroying the grass they graze on unless they take science more seriously (and not as a toddler idea of "generating and testing ideas", which is only a small part of the story).

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#26

To produce scientific work, one needs certain raw materials: 1. Data 2. Access to past works Once you have these, only then can discoveries can be made, and papers be written. How does this software get these? I am assuming they have to be provided up-front to the software for each job.

It will analyse what it is given, but it will not have the ability to say, "hang on, these results are interesting, I wonder what will happed if I pour a different liquid into the drum and spin it at the same speed?" LLMs, especially the latest ones, are decent at analysis of input, but disappointing at producing creative output. I am running a series of experiments using Gemini 1.5 and found it capable of producing…

Interestingly your comment is the very opposite of my experience with LLMs.

You can rely on them for creative stuff (write a poem, short story, etc.), but you cannot depend on them for factual stuff. Time and time again they will state "facts" that turn out to be false, and so now I no longer trust it for anything without manual verification. And since I then need to do the research myself for the verification, I rarely find LLMs helpful, except occasionally for initial exploration of some topic.

You used summarization as an example, but whether they are fundamentally good at that is even debatable, e.g. https://ea.rna.nl/2024/05/27/when-chatgpt-summarises-it-actu...

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#27
Some samples of the generated papers are in the SI of their paper. It’s be interesting if some of you ML guys dug into them. The fact that they built another AI system to review the papers seems really shaky, this is where human feedback would be most valuable.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#28

But OpenAI said LLMs can't innovate until human-level reasoning and long-term agenthood is solved. [1] Referring to their precious 5 stages to classify AI before it reaches the scary "beyond" levels of intelligence... presumably at that point they get the feds involved to reg cap the field, so genuine is the fear of the pace they've set. It's clear OpenAI is a hype company knocking over glass bottle stacks at its own…

sam made comments on Twitter we hit level 2 - we’ll know more in the coming weeks if he’s right.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#29
Exciting and very cool! I look forward to the continued improvement in this area. Especially when the loop is closed within Sakana and you can say "this discovery was made by The AI Scientist" as part of another paper.

If I might offer some small feedback on the blog post:

- Alt-text and/or caption of the initial image would be helpful for screen readers

- Using both "dramatically" and "radically" in one sentence to describe near future improvements seems a bit much.

- When talking about the models used, "Sonnet" could either be 3.0 Sonnet or 3.5 Sonnet and those have pretty different capabilities.

Thanks again for the impressive work!

Post reply on HN