> It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out. Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time... https://www.lesswrong.com/w/instrumental-convergence
> many possible Y-goals would concentrate probability into this X-strategy being used
The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…
Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is "aware", it's not trained to care, i.e. it's not an aligned model (alignment is considered hard).
Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.
> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…
i've always been under the assumption that "AI Safety" is baked into the training of the models and not a parameter that can be turned up or down. So if someone breaks into Anthropic one night and makes a full copy of Mythos or whatever then that model they copied is fully capable and not lobotomized? That raises questions because, if you believe all the PR, that's equivalent to breaking into a research university and stealing an entire bio/chem weapons research department.
edit: if the above is the case then we should just assume it's already happened because of the value to both goodguys(tm) and badguys(tm).
The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits m…
Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…
I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off.
OpenAI said:
> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
It was a test of raw capabilities of the underlying model.
Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.
There are many science fiction stories about amoral humans, cyborgs, and AIs. Paperclip maximization is a relatively recent meme, but the 'gray goo scenario' has been widely written about: https://en.wikipedia.org/wiki/Gray_goo
Call me obtuse but the way this is being portrayed by the companies involved, media, seems a little odd to me. Didn't the model + harness do what was asked? If I ask a coding agent to write a very clever piece of code and it turns out impressively clever, it did what I asked.
> Didn't the model + harness do what was asked? Depends exactly what they asked it to do, but it very clearly didn't do what was intended, or what an honest human would do. Stop trying to find a gotcha.
>Stop trying to find a gotcha.
Surely notorious liar Sam Altman wouldn’t lie this time.
Remember Stuxnet? Hack one of these, wait until the right compounds are physically loaded, then execute https://www.sigmaaldrich.com/US/en/products/chemistry-and-bi...
Stuxnet misconfigured industrial equipment that was already set up to run, and all it did was break that equipment. This scenario sets a much, much higher bar.
I guarantee there's misconfigured chemical analysis equipment out there exposed to the internet.
I don't think the grandparent was implying the AI would be controlling robot arms to mix things directly (or at least I didn't interpret as such), but it could very well sit in the network until it notices two dangerous compounds in the same machine, and trigger a breakage that causes a harmful mixture. Break a vial containing a virus, then break two more that cause an emergency evac and maybe that's enough to get something out there.
I think the mental model people have about this is that pre-AI there were humans picking individual targets and post-AI the computer itself randomly picks targets. But you get the same unexpected collateral damage outcome when a human misconfigures a decent pentest tool.
Also, that’s the wrong mental model. Anyone who runs a sass platform or a website knows that the Internet is already full of millions and millions and millions of bots and scripts and other random shit that’s always trying to attack you, often completely randomly. Security is always a battle between good and evil. All I can say is that if you’re in charge of keeping something secure, you should probably try to get yo…
That's exactly what they are doing. Out of all the civilian applications that AI capabilities have, why has this surfaced as a priority for demonstration? Probably because of money.