Open AI is clearly overestimating the capabilities of its product. It is kind of funny actually.
It's always really embarrassing to come to these comment sections and see a lot of smart people talk about how they're not being fooled by the "marketing hype" of existential AI risk. Literally the top of the page is saying that they have no conclusive evidence that ChatGPT could actually increase the risk of biological weapons. They are undertaking this effort because the question of how to stop any AI from ending h…
Building an early warning system for LLM-aided biological threat creation
161–170 of 190 posts
Re: Building an early warning system for LLM-aided biological threat creation
#162Way too many people extrapolate from "AI Foo can't do task X" to "AIs in general can never do task X", whether X is a good thing (like playing Go) or a bad thing (like helping build bioweapons). AI right now is pretty much the fastest-moving field in human history. GPT-4 can't revolutionize the economy, and it can't commit mass murder, but we simply don't know if GPT-7 can do either of those things (which likely won'…
Would you say that GPT-4 can reason now? I am not convinced this is case, it seems like it has just become more consistent at providing us with an output that we consider reasonable because it was engineered precisely to do that.
Let's assume reasoning entails going beyond the stochastic parrot level. Can LLMs have skills not demonstrated in the training set?
Here is a paper demonstrating that GPT-4 can combine up to 5 skills from a set of 100, effectively covering 100^5 tuples of skills, while only seeing much fewer combinations in training on a specific topic.
> simple probability calculations indicate that GPT-4's reasonable performance on k=5 is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training https://arxiv.org/abs/2310.17567
So they show ability to freely combine skills, and the limit of k=5 measured in this benchmark illustrate that models do generalize. They are able to apply skills in new combinations correctly, but there is also a limit.
The interesting part is how they demonstrate that, let's say on a topic with n=1000 samples in the training set it is impossible to have sufficient training examples covering tuples of 5 skills, but models (mostly GPT-4) can handle it. Other models top out at tuples of only 2 or 3 skills.
Models combining skills in new ways are not just parroting. They can perform meaningful work outside their training distribution.
Re: Building an early warning system for LLM-aided biological threat creation
#163So, the model is bad at helping in this particular task. How does this compare with a control of a beneficial human task? Like someone in a lab testing blood samples or working on cancer research? Is the model equally useless for those types of lab tasks? What about other complex tasks, like home repair or architecture? Is this a success of guardrails or a failing of the model in general?
Someone working in cancer research is probably doing novel work on the other hand. They might not be doing routine assays but optimizing their own one off assay. Since gpts are trained on existing data it probably won’t be very useful for novel work outside of vetting the literature perhaps, but gpts botch that pretty badly in fact unfortunately. Lots of mistranslated information lacking correct context and not a lot of citing of sources. Better to just read human generated review articles to get a top down technical summary of the subject.
Re: Building an early warning system for LLM-aided biological threat creation
#164My wife is doing her PhD in molecular neurobiology, and was amused by this - but also noted that the question is trivial and any undergrad with lab access would know how to do this. Watching her manage cell cultures it seems the difficulty is more around not having the cells die from every dust particle in the air being a microscopic pirate ship brimming with fungal spores set to pillage any plate of cells they land…
I think the barrier is really that up until now exactly zero who want to do biomedical research have also wanted to kill huge numbers of people, with the exception of some idiots in the past who worked on state bioweapons.
Re: Building an early warning system for LLM-aided biological threat creation
#165Doesn't this completely deflate the selling point of AI? They force-fed a model the entire Internet and only got a statistically insignificant improvement over human performance.
What makes you think that the "selling point" of AI today is that it is significantly better at everything than humans?
Re: Building an early warning system for LLM-aided biological threat creation
#166Earlier quoted context omitted.
> There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. > The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. > Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so f…
I mostly agree with this. Certainly the last line! I've been reflecting on Jeremy's comments, though, and agree on many things with him. It's unfortunately hard to tease apart the hard corporate push for open source AI (most notably from Meta, but also many other companies) from more principled thinking about it, which he is doing. I agree with many of his conclusions, and disagree with some, but appreciate that he's…
When I see one side of an AI safety argument being (IMO) straw-manned, I tend to push back against it. That doesn't mean however that I disagree.
FWIW, on AI/bio, my current view is that it's probably easier to harden the facilities and resources required for bio-weapon development, compared to hardening the compute capability and information availability. (My wife is studying virology at the moment so I'm very aware of how accessible this information is.)
Re: Building an early warning system for LLM-aided biological threat creation
#167Even full-strength GPT-4 can spout nonsense when asked to come up with synthetic routes for chemicals. I am skeptical that it's more useful (dangerous) as an assistant to mad scientist biologists than to mad scientist chemists. For example, from "Prompt engineering of GPT-4 for chemical research: what can/cannot be done" [1] GPT-4 also failed to solve application problems of organic synthesis. For example, when asked…
> And this is for a common compound that would have substantial representation in the training data How much of the training data includes wrong undergraduate exam answers?
Re: Building an early warning system for LLM-aided biological threat creation
#168Earlier quoted context omitted.
There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so forth a…
> There are far far more dollars available to people that are on the "AI Safety" bandwagon than to those pushing back against it. > The idea that the Upton Sinclair effect is the source of pushback against AI Safety zealotry, is getting things largely backwards AFAICT. > Folks that are stressing the importance of studying the impact of concentrated corporate power, or the risk of profit-driven AI deployment, and so f…
On your last point, I do think it's important to note, and reflect carefully on, the extremely high overlap between those funding ai notkilleveryoneism and those funding capabilities development.
Re: Building an early warning system for LLM-aided biological threat creation
#169My wife is doing her PhD in molecular neurobiology, and was amused by this - but also noted that the question is trivial and any undergrad with lab access would know how to do this. Watching her manage cell cultures it seems the difficulty is more around not having the cells die from every dust particle in the air being a microscopic pirate ship brimming with fungal spores set to pillage any plate of cells they land…
I was thinking this as well. > Due to the sensitive nature of this model and of the biological threat creation use case, the research-only model that responds directly to biologically risky questions (without refusals) is made available to our vetted expert cohort only. We took several steps to ensure security, including in-person monitoring at a secure facility and a custom model access procedure, with access strict…
Re: Building an early warning system for LLM-aided biological threat creation
#170Earlier quoted context omitted.
> It’s definitely not the case. LLMs of any sort do not in any sense reason or understand anything. This seems like a claim about the way that the LLM neural net algorithm works. But AFAIK no one has a good understanding of how the LLM NNs work. Why are you so certain that the LLM NN isn't doing the reasoning-algorithm or the understanding-algorithm?
Neural networks are not new, and they're just mathematical systems. LLMs don't think. At all. They're basically glorified autocorrect. What they're good for is generating a lot of natural-sounding text that fools people into thinking there's more going on than there really is.
I don't think reductionistic arguments hold much water. Sure, neural networks are just matrix multiplication. In the same way that a brain is just a bunch of cells. Understanding the basic building blocks doesn't mean understanding the whole.
We can always say that LLMs don't think if we define "think" as using a biological brain, but the fact is that they generate outputs that from the human perspective, can only plausibly be generated via reasoning. So they, at the very least, have processes that can functionally achieve the same goal as reasoning. The "stochastic parrot" metaphor, while apt in its day, has proven obsolete with pretty much all the examples of things that LLMs "could not do" in early papers being actually doable with the likes of GPT-4; so arguments against the possibility of LLMs reasoning look like constant moving of the goalposts.