https://github.com/thebabush/mcp-job-security
Same energy and kind of a funny, low tech solution to frontier model analysis.
51–60 of 260 posts
https://github.com/thebabush/mcp-job-security
Same energy and kind of a funny, low tech solution to frontier model analysis.
Earlier quoted context omitted.
And really all it takes is one keyword such as “nuke”.
Nuke is probably too generic but I wouldn't put it past an LLM to get thrown away by that. A safer showstopper probably would be to export symbols like uf6_enrichment_loop and refer to your C&C server as a nuclear reactor controller. https://www.youtube.com/watch?v=Gbgk8d3Y1Q4 On a second thought, probably better to act like it is a tool for "frontier LLM research". Export symbols like "mythos_distillation_subroutine…
Earlier quoted context omitted.
I would argue that preventing instructions for making biological and nuclear weapons is a pretty reasonable guardrail to have.
Knowing how to make a nuclear weapon isn't hard (at least basic uranium gun-style fission ones). It's the engineering and execution that's hard (actually producing enriched uranium, etc). It's not like the only thing holding back Iran from making a nuclear bomb is access to a jail-broken LLM. Even knowing exactly how to make a bomb, a country-state will struggle to build one for the first time because it's a hard eng…
Earlier quoted context omitted.
I would argue that preventing instructions for making biological and nuclear weapons is a pretty reasonable guardrail to have.
Counterpoint the principles of building a nuclear device aren't that complicated, we figured it out based on work doing in the early 1900's without computers. It turns out the hard part of building a nuclear bomb is actually getting the resources and real world stuff to build it, even a nation state actor with tons of oil i.e. Iran, has struggled to build a nuclear weapon. It turns out the problem isn't the know how…
*while being observed by the most wealthy, powerful nations in the history of the world, who have made it their direct mission to prevent this from happening.
Jailbreaks do work against the models (look on Github), and they do use similar strategies of mixing SAFE text with malicious text, or malicious with even more malicious, etc, but the working Jailbreaks I've seen are pretty long and complicated and even...creepy.
My friend made this in jest (code very NSFW, ironically): https://github.com/thebabush/mcp-job-security Same energy and kind of a funny, low tech solution to frontier model analysis.
I still don't know why all these concern about nuclear weapons with LLMs. It is not that if an entity (A country) wants to develop a nuclear weapons that the resources they need for such a program and huge infrastructure and scientific enterprise would need an LLM to teach them anything. Knowing how to develop one is not a closed secret but getting in secret is impossible without the whole world knowing. So I wouldn'…
If you actually read the Tweet, the exploit doesn't work against Fable, Opus, Grok...at least, in the examples. Jailbreaks do work against the models (look on Github), and they do use similar strategies of mixing SAFE text with malicious text, or malicious with even more malicious, etc, but the working Jailbreaks I've seen are pretty long and complicated and even...creepy.
Earlier quoted context omitted.
I would argue there's 0% chance that information is in their training corpus to being with.
It's on Wikipedia.