Live data from Hacker News

The AI Scientist: Towards Automated Open-Ended Scientific Discovery

sakana.ai

61–70 of 144 posts

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#61
As someone 'in academia', I worry that tools like this fundamentally discard significant fractions of both the scientific process and why the process is structured that way.

The reason that we do research is not simply so that we can produce papers and hence amass knowledge in an abstract sense. A huge part of the academic world is training and building up hands-on institutional knowledge within the population so that we can expand the discovery space.

If I went back to cavemen and handed them a copy of _University Physics_, they wouldn't know what to do with it. Hell, if I went back to Isaac Newton, he would struggle. Never mind your average physicist in the 1600s! Both the community as a whole, and the people within it, don't learn by simply reading papers. We learn by building things, running our own experiments, figuring out how other context fits in, and discussing with colleagues. This is why it takes ~1/8th of a lifetime to go from the 'world standard' of knowledge (~high school education) to being a PhD.

I suppose the claim here is that, well, we can just replace all of those humans with AI (or 'augment' them), but there are two problems:

a) the current suite of models is nowhere near sophisticated enough to do that, and their architecture makes extracting novel ideas either very difficult or impossible, depending on who you ask, and;

b) every use-case of 'AI' in science that I have seen also removes that hands-on training and experience (e.g. Copilot, in my experience, leads to lower levels of understanding. If I can just tab-complete my N-body code, did I really gain the knowledge of building it?)

This is all without mentioning the fact that the papers that the model seems to have generated are garbage. As an editor of a journal, I would likely desk-reject them. As a reviewer, I would reject them. They contain very limited novel knowledge and, as expected, extremely limited citation to associated works.

This project is cool on its face, but I must be missing something here as I don't really see the point in it.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#62
post #51

Earlier quoted context omitted.

Only the last one is in any way actually bad and even then it should be in the interest of the company using it to fix it promptly.

Deaths in car crashes and copyright laundering by big corporations are not bad in any way at all?

I would say that car crashes are bad, even though they already happen and the motivation behind AI is to reduce them by being less bad than a human.

I think it is a mistake to trust 1st party statistics on the quality of the AI, the lack of licence for level 5 suggests the US government is unsatisfied with the quality as well, but in principle this should be a benefit. When it actually works.

Copyright is an appalling mess, has been my whole life. But no, the economic threat to small copyright holders, individual artists and musicians, is already present by virtue of a globalised economy massively increasing competition combined with the fact the resulting artefacts can be trivially reproduced. What AI does here needs consideration, but I have yet to be convinced by an argument that what it does in this case is bad.

All these things will likely see a return to/increase in patronage, at least for those arts where the point is to show off your wealth/taste; the alternative being where people just want nice stuff, for which mass production has led to the same argument since Jaquard was finding his looms smashed by artisans who feared for their income.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#63
I'm not a scientist at all, but I am often involved in hand-holding scientists when it comes to dealing with computers.

My impression so far is that science is plagued with deliberate and accidental fraud when it comes to data collection and cataloguing. Also, this is a spectrum, not two distinct things. I often see researchers simply unwilling to do the right thing to verify that the data collected are correct and meaningful as soon as "workable" results can be produced from the data. Some will go further and mess with the data to make results more "workable" though...

Second problem is understanding the data. Often times it happens that people who end up doing research don't quite understand the subject matter of the research. This is especially popular with medicine, where it's overwhelmingly common for eg. research into various imaging modalities to be done by computer scientists who couldn't find a liver cancer the size of a coconut in the sharpest textbook abdominal image.

My impression is also that by far these two problems outweigh the problems that could potentially be solved by adding AI into the mix. These are the systemic organization problems of perverse incentives and vicious practices, and no amount of AI is going to do anything about it... because it's about people. People's salaries, careers, friendships etc.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#65

Earlier quoted context omitted.

It would also be sad to see the scientific system destroyed by a wave of automatically generated papers that no human has the capacity to verify. It's not hard to generate ideas, it's hard to generate reliable and relevant ideas. Such AI science generators are destroying the grass they graze on unless they take science more seriously (and not as a toddler idea of "generating and testing ideas", which is only a small…

we already have a wave of papers that no human has the capacity to verify

I have also, unless I hallucinated it, read accusations on this very site of peer reviewed papers that were at least partly generated by LLMs.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#66
post #55
post #54

Earlier quoted context omitted.

It's quite common for new species to kill off old species. We ourselves have obliterated many species that we outcompeted for resources.

As if software is the same thing as a new biological species. I am just so bored of reading bullshit like this. If you really believe this then you need to level up your level of education and learning. It is not good.

> If you really believe this then you need to level up your level of education and learning. It is not good.

How does your level of education and learning compare to Nobel Prize winner Dr. Geoffrey Hinton (father of deep learning, 10%-50% chance that AI will kill everyone), Dr. Dan Hendrycks (GELU inventor, >80%), Dr. Jan Leike (DeepMind, OpenAI, 10%-90%), Dr. Paul Christiano (OpenAI, Time 100 2023, UK Frontier Taskforce advisory board, 46%), etc.?

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#67
post #53
post #27

Some samples of the generated papers are in the SI of their paper. It’s be interesting if some of you ML guys dug into them. The fact that they built another AI system to review the papers seems really shaky, this is where human feedback would be most valuable.

Tried reading the 'low-dimensional diffusion' one. Not an expert on diffusion by any means, but the very premise of the paper seems like bullshit. It claims that 'while diffusion works in high-dimensional datasets, it struggle in low-dimensional settings', which just makes no sense to me? Modeling high-dimensional data is just strictly harder than low-dimensional one. Then when you read the intro, it's full of 'blank…

I would agree with your analysis.

Note that it cites TabDDPM in the related work, but that is for diffusione on tabular data! While most tabular data is low-dimensional, the type of low-dimensional data tackled in the paper is not tabular!

I'm also not quite sure how the linear upscaling is supposed to help, as it can be absorbed into the first layer of the following MLP, so I would rather think that the performance improvement (if any, the numbers are quite close and lack standard errors) is either due to the increased number of trainable parameters or some kind of ensembling effect (essentially the mixture of experts point made by the human authors).

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#69
> The AI Scientist is designed to be compute efficient. Each idea is implemented and developed into a full paper at a cost of approximately $15 per paper. While there are still occasional flaws in the papers produced by this first version (discussed below and in the report), this cost and the promise the system shows so far illustrate the potential of The AI Scientist to democratize research and significantly accelerate scientific progress.

This is a particularly confusing argument in my opinion. Is the underlying assumption that everyone wants, or even needs, white papers that they can claim they created?

Let's just assume this system actually works and produces high quality, rigorous research findings. Reducing that process down to a dollar amount and driving that cost to near zero doesn't democratize anything, it cheapens it to the point of being worthless.

This honestly reads more as a joke article trolling today's academic process and the whole publish or perish mentality. From that angle, the article is a success in my book. As an announcement for a new ML tool though, I just don't get it.

Re: The AI Scientist: Towards Automated Open-Ended Scientific Discovery

#70
post #61

As someone 'in academia', I worry that tools like this fundamentally discard significant fractions of both the scientific process and why the process is structured that way. The reason that we do research is not simply so that we can produce papers and hence amass knowledge in an abstract sense. A huge part of the academic world is training and building up hands-on institutional knowledge within the population so tha…

We already substitute "good authority" (be it consensus or a talking head) for "empirical grounding" all the time. Faith in AI scientific overlords seems a trivial step from there.
Post reply on HN