Live data from Hacker News

Building an early warning system for LLM-aided biological threat creation

openai.com

161–170 of 190 posts

Re: Building an early warning system for LLM-aided biological threat creation

#161
post #48

Open AI is clearly overestimating the capabilities of its product. It is kind of funny actually.

It's always really embarrassing to come to these comment sections and see a lot of smart people talk about how they're not being fooled by the "marketing hype" of existential AI risk. Literally the top of the page is saying that they have no conclusive evidence that ChatGPT could actually increase the risk of biological weapons. They are undertaking this effort because the question of how to stop any AI from ending h…

The type of ai that can end humanity will not be a chatbot. Knowing statistically which word is liable to be next in a sentence is what these tools actually do when you take off the marketing glowup. Whats more likely is that these sorts of ais take a lot of jobs that produce churn content, a good thing if people spent that waste of a time doing something creative instead perhaps, but unfortunately our world does not alleviate us from busy work to provide us with food shelter and creative outlets in its wake. Quite the opposite. Maybe thats what needs our attention in a more automated world than anything from sci fi.

Re: Building an early warning system for LLM-aided biological threat creation

#162

Way too many people extrapolate from "AI Foo can't do task X" to "AIs in general can never do task X", whether X is a good thing (like playing Go) or a bad thing (like helping build bioweapons). AI right now is pretty much the fastest-moving field in human history. GPT-4 can't revolutionize the economy, and it can't commit mass murder, but we simply don't know if GPT-7 can do either of those things (which likely won'…

Would you say that GPT-4 can reason now? I am not convinced this is case, it seems like it has just become more consistent at providing us with an output that we consider reasonable because it was engineered precisely to do that.

> Would you say that GPT-4 can reason now?

Let's assume reasoning entails going beyond the stochastic parrot level. Can LLMs have skills not demonstrated in the training set?

Here is a paper demonstrating that GPT-4 can combine up to 5 skills from a set of 100, effectively covering 100^5 tuples of skills, while only seeing much fewer combinations in training on a specific topic.

> simple probability calculations indicate that GPT-4's reasonable performance on k=5 is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training https://arxiv.org/abs/2310.17567

So they show ability to freely combine skills, and the limit of k=5 measured in this benchmark illustrate that models do generalize. They are able to apply skills in new combinations correctly, but there is also a limit.

The interesting part is how they demonstrate that, let's say on a topic with n=1000 samples in the training set it is impossible to have sufficient training examples covering tuples of 5 skills, but models (mostly GPT-4) can handle it. Other models top out at tuples of only 2 or 3 skills.

Models combining skills in new ways are not just parroting. They can perform meaningful work outside their training distribution.

Re: Building an early warning system for LLM-aided biological threat creation

#163
post #9

So, the model is bad at helping in this particular task. How does this compare with a control of a beneficial human task? Like someone in a lab testing blood samples or working on cancer research? Is the model equally useless for those types of lab tasks? What about other complex tasks, like home repair or architecture? Is this a success of guardrails or a failing of the model in general?

Someone working in a lab doing routine blood work isn’t going to benefit from this. They aren’t doing anything novel just running the same assay a hundred times a week. A machine can do that job without ai today.

Someone working in cancer research is probably doing novel work on the other hand. They might not be doing routine assays but optimizing their own one off assay. Since gpts are trained on existing data it probably won’t be very useful for novel work outside of vetting the literature perhaps, but gpts botch that pretty badly in fact unfortunately. Lots of mistranslated information lacking correct context and not a lot of citing of sources. Better to just read human generated review articles to get a top down technical summary of the subject.

Re: Building an early warning system for LLM-aided biological threat creation

#164

My wife is doing her PhD in molecular neurobiology, and was amused by this - but also noted that the question is trivial and any undergrad with lab access would know how to do this. Watching her manage cell cultures it seems the difficulty is more around not having the cells die from every dust particle in the air being a microscopic pirate ship brimming with fungal spores set to pillage any plate of cells they land…

When we did this kind of thing in high school we had huge problems with contamination, but I don't agree that it is so hard.

I think the barrier is really that up until now exactly zero who want to do biomedical research have also wanted to kill huge numbers of people, with the exception of some idiots in the past who worked on state bioweapons.

Re: Building an early warning system for LLM-aided biological threat creation

#165

Doesn't this completely deflate the selling point of AI? They force-fed a model the entire Internet and only got a statistically insignificant improvement over human performance.

What makes you think that the "selling point" of AI today is that it is significantly better at everything than humans?

Thats openais marketing stance honestly. Thats how they define general ai actually, on economic vs technical terms.

Re: Building an early warning system for LLM-aided biological threat creation

#166

Earlier quoted context omitted.

> There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. > The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. > Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so f…

I mostly agree with this. Certainly the last line! I've been reflecting on Jeremy's comments, though, and agree on many things with him. It's unfortunately hard to tease apart the hard corporate push for open source AI (most notably from Meta, but also many other companies) from more principled thinking about it, which he is doing. I agree with many of his conclusions, and disagree with some, but appreciate that he's…

Thank you Michael. I'm not even sure I disagree with you on many things -- I think things are very complicated and nuanced and am skeptical of people that hold overly strong opinions about such things, so I try not to be such a person myself!

When I see one side of an AI safety argument being (IMO) straw-manned, I tend to push back against it. That doesn't mean however that I disagree.

FWIW, on AI/bio, my current view is that it's probably easier to harden the facilities and resources required for bio-weapon development, compared to hardening the compute capability and information availability. (My wife is studying virology at the moment so I'm very aware of how accessible this information is.)

Re: Building an early warning system for LLM-aided biological threat creation

#167

Even full-strength GPT-4 can spout nonsense when asked to come up with synthetic routes for chemicals. I am skeptical that it's more useful (dangerous) as an assistant to mad scientist biologists than to mad scientist chemists. For example, from "Prompt engineering of GPT-4 for chemical research: what can/cannot be done" [1] GPT-4 also failed to solve application problems of organic synthesis. For example, when asked…

> And this is for a common compound that would have substantial representation in the training data How much of the training data includes wrong undergraduate exam answers?

God help us all if it crawled Chegg

Re: Building an early warning system for LLM-aided biological threat creation

#168
post #132

Earlier quoted context omitted.

There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so forth a…

> There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. > The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. > Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so f…

I'm not really trying to rebut Michael's argument -- I think it's true, to an extent, some of the time. But I think it's more true more of the time in the reverse direction. So I don't think it's a good argument. And more importantly, I think it fails to properly grapple with the ideas, instead using an ad hominem approach to discarding them somewhat thoughtless.

On your last point, I do think it's important to note, and reflect carefully on, the extremely high overlap between those funding ai notkilleveryoneism and those funding capabilities development.

Re: Building an early warning system for LLM-aided biological threat creation

#169
post #156

My wife is doing her PhD in molecular neurobiology, and was amused by this - but also noted that the question is trivial and any undergrad with lab access would know how to do this. Watching her manage cell cultures it seems the difficulty is more around not having the cells die from every dust particle in the air being a microscopic pirate ship brimming with fungal spores set to pillage any plate of cells they land…

I was thinking this as well. > Due to the sensitive nature of this model and of the biological threat creation use case, the research-only model that responds directly to biologically risky questions (without refusals) is made available to our vetted expert cohort only. We took several steps to ensure security, including in-person monitoring at a secure facility and a custom model access procedure, with access strict…

Truth. AI researchers need to work with actual scientists or they ended up with the same brain-in-a-jar problems that their models have...

Re: Building an early warning system for LLM-aided biological threat creation

#170

Earlier quoted context omitted.

> It’s definitely not the case. LLMs of any sort do not in any sense reason or understand anything. This seems like a claim about the way that the LLM neural net algorithm works. But AFAIK no one has a good understanding of how the LLM NNs work. Why are you so certain that the LLM NN isn't doing the reasoning-algorithm or the understanding-algorithm?

Neural networks are not new, and they're just mathematical systems. LLMs don't think. At all. They're basically glorified autocorrect. What they're good for is generating a lot of natural-sounding text that fools people into thinking there's more going on than there really is.

Not new, but we don't understand how they work at the large scale.

I don't think reductionistic arguments hold much water. Sure, neural networks are just matrix multiplication. In the same way that a brain is just a bunch of cells. Understanding the basic building blocks doesn't mean understanding the whole.

We can always say that LLMs don't think if we define "think" as using a biological brain, but the fact is that they generate outputs that from the human perspective, can only plausibly be generated via reasoning. So they, at the very least, have processes that can functionally achieve the same goal as reasoning. The "stochastic parrot" metaphor, while apt in its day, has proven obsolete with pretty much all the examples of things that LLMs "could not do" in early papers being actually doable with the likes of GPT-4; so arguments against the possibility of LLMs reasoning look like constant moving of the goalposts.

Post reply on HN