Live data from Hacker News

Claude Science

claude.com

91–100 of 199 posts

Re: Claude Science

#91
post #44

I built one of the connected tools included in this launch (the Biomni HPC [1]), and I have spent an inordinate amount of my life working on this problem. (I also worked at Anthropic, but not on this product.) As other comments have pointed out, this is for data science – but it's capable of more than making plots and writing papers [2]. It has integrations with many databases and computational tools, including a res…

How do you validate this kind of work to weed out any confabulating by the LLMs?

Re: Claude Science

#92
A "standing review agent" seems to be one of the main differences beyond the new connectors and in place visualisation tools.

>A standing reviewer agent. This runs in the background during a session, checking citations against sources, flagging numbers it can't trace back to evidence, and catching figures that don't match the code that supposedly generated them. That's not something Code or Cowork do automatically — you'd have to ask Claude to double-check itself as a separate step.

Re: Claude Science

#93
post #88
post #84

Earlier quoted context omitted.

Just as a counterpoint ML and AI research has become much more reproducible over time. I feel like this is relevant because ML / AI researchers are huge power users of AI tools. Between 2016 and 2021 the share of ML/ robotics/ AI researchers being reproducible (ie contianing code and similar instructions to reproduce) doubled [1]. The major US labs have gone largely closed source (I.e. they no longer publish frontier…

That's good to hear about ML and AI research, but most research is not based on computers and so would require laboratory setups to reproduce. Not only is trying to reproduce such findings (beyond what is effectively a sanity check) through simulations a lost cause, if AI can reproduce such research it would be capable of doing such research itself... in which case it would be far more fruitful to use AI to do furthe…

This thread is about a product based on fully reproducible research though so I feel like we should stay grounded. Claude science is meant to be used in the context of reproducible science research, there is a decent reason to not be cynical on future research being reproducible.

> if AI can reproduce such research it would be capable of doing such research itself

Well there is a big distinction between research validation and research generation, it is generally much easier to verify that a math proof is true or false than to find a truly novel proof.

But yes in the long run I’d think AI will be doing tons of research and it will by default reproducible. So maybe we’re aligned after all?

Re: Claude Science

#94
post #5

Science isn’t suffering from a lack of papers. It’s suffering from a lack of good papers. Making it easier to just pump out paper-mill publications is about the last thing science needs right now.

Scientific research is suffering from a reproducibility crisis. Not a publication crisis. LLM's aren't going to solve reproducibility issues.

Underlying reproducibility is integrity.

Underlying integrity is rigor.

Underlying rigor is education.

It goes deep, for sure, IMO.

Re: Claude Science

#97

The fact that we are coming up on a month of Fable being unavailable with essentially zero actual signal from Anthropic around when it may be back is crazy to me. Yet still we have these random new products coming out?

This thing is also surprising considering Fable was not allowed to answer any biology questions.

Re: Claude Science

#98
post #48

Earlier quoted context omitted.

My take based on the video is that they're thinking more about bioinformatics, which might technically fall under the "data science" umbrella depending how you define your terms, but which is not described that way in common usage. It's the content that determines the sort of science, not the toolchain.

Honestly quite excited to see what can happen here, I think biology has generally had a lack of data science expertise.

Tell us, what gives you that impression?

Re: Claude Science

#100
Tried this to see how it goes in my particular field - computational design of RNAi-based biopesticides. One-shotted a design for targeting the DvSnf7 transcript of western corn rootworm. It took a fairly naive approach (maybe how a 1st year PhD student would go about it), but got the job done. Also noted caveats with its approach (e.g. using mammalian design rules, limited off-target screening). Not bad really. But also not great. When its flaws were pointed out, the AI determined that it could have taken a more informed approach. Then Opus 4.8's safety system flagged the session.
Post reply on HN