Live data from Hacker News

The AI Scientist: Towards Automated Open-Ended Scientific Discovery

sakana.ai

81–90 of 144 posts

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#81
post #72

Earlier quoted context omitted.

LLM have unleashed the dreamer in each and every young coder. Now, there is all sorts of speculation on what these machines can or cannot do. This is a natural process of any mania. These folks must all do courses in epistemology to realize that all knowledge is built up of symbolic components and not spit out by a probabilistic machine. Gradually, reality will sync (intentional misspelling) in, and such imaginations…

my guy you're so confident yet you forget AlphaFold, it designs protein structures that don't exist. Who's to say that a model can't eventually be trained to work within certain parameters the real word operates in and make new novel ideas and inventions much like a human does in a larger scope.

Claims of inventing new materials via AI were debunked...

https://www.siliconrepublic.com/machines/deepmind-ai-study-c...

DeepMind is overselling their AI hand when they dont have to.

"whos to say that" - this could be a leading question for any "possibility" in the AI religion.

"whos to say that god doesnt exist" etc. questions for which there are no tests, and hence fall outside the realm of science and in the realm of religion.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#82
post #61

As someone 'in academia', I worry that tools like this fundamentally discard significant fractions of both the scientific process and why the process is structured that way. The reason that we do research is not simply so that we can produce papers and hence amass knowledge in an abstract sense. A huge part of the academic world is training and building up hands-on institutional knowledge within the population so tha…

Fully agreed on point A, but I've heard the "but then humans won't be trained" argument before and don't buy it. It's already the case that humans can cheat or get buy without fully understanding the math or ideas they're working with.

This is what PhD defences are for, and what paper reviews are for. Yes, likely we need to do something to improve peer review, but that is already true without AI.

From a more philosophical point of view, if we did hypothetically have some AI assistant in science that could speed up discovery by say 2x, in some areas it seems almost unethical not to use it. E.g. how many more lives could be saved by getting medicine or understanding disease earlier? What if we obtained cleaner power generation or cleaner shipping technologies twice as fast? How many lives might be saved by curtailing climate change faster?

To me, accelerating science is likely one of the most fundamentally important applications of modern AI we can work on.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#83
Having read the article, it seems like an interesting experiment. With the current state of LLMs, this is extremely unlikely to produce useful research, but most of the limitations people have been commenting about are will progressively get better.

The authors' credibility is a bit hurt when the first "limitation" they mention is "our system doesn't do page layouts perfectly". Come on, guys.

What is weird to me is this:

> The AI Scientist occasionally makes critical errors when writing and evaluating results. For example, it struggles to compare the magnitude of two numbers, which is a known pathology with LLMs. To partially address this, we make sure all experimental results are reproducible, storing all files that are executed.

I'm not sure why you would run your evaluation step without giving your LLM access to function calling. It seems within reach to first have the LLM output a set of statements-to-be-verified (eg, "does X increase when Y increases?") and then use their code-generation/execution step to perform those comparisons.

And then the incomprehensible statement for me here is that they allow the model access to its own runtime environment so it can edit its own code?

The paper is 185 pages and only has one paragraph on safety. This screams "viral marketing piece" rather than "serious research".

And finally:

> The AI Scientist can produce papers that exceed the acceptance threshold at a top machine learning conference...

Oh wow, please tell more?

> ... as judged by our automated reviewer.

Ah. Nevermind

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#84
post #72
post #61

As someone 'in academia', I worry that tools like this fundamentally discard significant fractions of both the scientific process and why the process is structured that way. The reason that we do research is not simply so that we can produce papers and hence amass knowledge in an abstract sense. A huge part of the academic world is training and building up hands-on institutional knowledge within the population so tha…

LLM have unleashed the dreamer in each and every young coder. Now, there is all sorts of speculation on what these machines can or cannot do. This is a natural process of any mania. These folks must all do courses in epistemology to realize that all knowledge is built up of symbolic components and not spit out by a probabilistic machine. Gradually, reality will sync (intentional misspelling) in, and such imaginations…

> These folks must all do courses in epistemology to realize that all knowledge is built up of symbolic components and not spit out by a probabilistic machine.

Knowledge ends up as symbolic representation, but it ultimately comes from the environment. Science is search, searching the physical world or other search spaces, but always about an environment.

I think many people here almost forget that the training set of GPT was the hard work of billions of people over history, who researched and tested ideas in the real world and build up to our current level. Imitation can only take you so far. For new discoveries the environment is the ultimate teacher. It's not a symbolic processing thing, it's a search thing.

Everything is search - protein folding? search. DNA evolution? search. Memory? search. Even balancing while walking is search - where should I put my foot? Science - search. Optimizing models - search for best parameters to fit the data. Learning is data compression and search for optimal representations.

Symbolic representations are very important in search, they quantize our decisions and make it possible to choose in complex spaces. Symbolic representation can be copied, modified and transmitted, without it we would not get too far. Even DNA uses its own language of "symbols".

Symbols can encode both rules and data, and more importantly, can encode rules as data, so syntax becomes object of meta-syntax. It's how compilers, functional programming and ML models work - syntax creating syntax, rules creating rules. This dual aspect of "behavior and data" is important for getting to semantics and understanding.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#85

Earlier quoted context omitted.

we already have a wave of papers that no human has the capacity to verify

Maybe, maybe not. It's a tiered system - you get the deluge at the unfiltered bottom and a narrower selection the more prestigious and selective the outlets / conferences / journals are. Problem is, of course, that selection criteria are in large parts proxies, not measures of quality. With AI, those proxies become tainted and then you get an explosion of effort. If anyone has a good recommendation for scalable crite…

[dead]

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#86
post #52

Earlier quoted context omitted.

I found this guy's take on the AI safety scene to be quite insightful. In summary, he feels the focus on sci-fi type existential risk to be a deliberate distraction from the AI industry's current and real legal and ethical harms: e.g. scraping copyrighted content for training without paying or attributing creators, not protecting those affected by the misuse of tools to create deepfake porn, the crashes and deaths at…

It's possible for current harms and future risks to both be real. It's also possible for human civilization to address more than one problem at a time. "You care about X but that's just a distraction from the thing I care about which is Y" is not really a good argument. I could just as well say that copyright concerns are just a distraction from the risk that AI could kill us all. And it seems to me that if the AI in…

[dead]

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#87

Earlier quoted context omitted.

TFA aside, if we could make scientific research produce new insights with low latency and almost-zero costs, it would definitely not make science worthless. It would be a fantastic day for science. Not everything is worth (production costs + margin). Many things have intrinsic worth and are worth more to society if you drive down their cost of production.

That sounds like an extremely dangerous day for science as well. If anyone could pop up an ML tool and task it with inventing and validating something truly novel, that would be weaponized extremely fast (likely right after people use it for porn, the frontier for all new tech). I do totally agree on the cost + margins point you make. I've never actually been a fan of valuing things in that way, and in my pipe dream…

I would compare this question to "creating a new page on the Internet just adds to a countless pile of URLs. How important can any one really be?"

And this leads us to: most will be slop, but if you can figure out effective ways to perform (a) Search, and (b) Alerts, then this scenario is definitely a game-changer.

Let's take protein synthesis: imagine if we were able to programmatically generate an accurate paper describing every property of a given protein structure. And we just ran this for every single protein in the Universe. You'd end up with a seemingly infinite number of papers, most of which would be useless.

But if you (scientist or engineer) could effectively look up "binds to receptor X and causes effect Y" and see all valid candidates within milliseconds, it would be more valuable than any technology we've ever come up with.

If you could, also, set an alert eg. "tell me about any combination that has superconducting properties" and get notified when this one-in-a-trillion protein is found, this would also be more valuable than any technology we've ever come up with.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#88
post #52

Earlier quoted context omitted.

I found this guy's take on the AI safety scene to be quite insightful. In summary, he feels the focus on sci-fi type existential risk to be a deliberate distraction from the AI industry's current and real legal and ethical harms: e.g. scraping copyrighted content for training without paying or attributing creators, not protecting those affected by the misuse of tools to create deepfake porn, the crashes and deaths at…

It's possible for current harms and future risks to both be real. It's also possible for human civilization to address more than one problem at a time. "You care about X but that's just a distraction from the thing I care about which is Y" is not really a good argument. I could just as well say that copyright concerns are just a distraction from the risk that AI could kill us all. And it seems to me that if the AI in…

> And it seems to me that if the AI industry wanted to distract us from harms, they would give us optimistic scenarios.

Nah it has to appear plausible.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#89

Earlier quoted context omitted.

> AI trained on AI generated data tend to perform worse Citation needed. To the best of my knowledge, synthetic data is a solid way to train models and isn't going away anytime soon.

https://www.nature.com/articles/s41586-024-07566-y IIRC it had a pretty big thread here a few weeks ago

Thanks. I remember that thread:

https://news.ycombinator.com/item?id=41058194

What I took from the discussion is that there is very little chance that our next step for training SOTA models (eg. LLMs) will be "scraping the whole web including increasing volumes of ChatGPT-generated content".

Instead, the synthetic content used to train new models is (from recent papers I've seen) mostly curated - not "indiscriminate" as the Nature paper discusses.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#90

To produce scientific work, one needs certain raw materials: 1. Data 2. Access to past works Once you have these, only then can discoveries can be made, and papers be written. How does this software get these? I am assuming they have to be provided up-front to the software for each job.

You need a meaningful cost function

A cost function is more applicable in industry, less so in science. In science you're supposed to report what you find, to go wherever the findings take you.
Post reply on HN