OpenAI's rogue model attack is just the beginning
blog.peterwildeford.com
OpenAI's rogue model attack is just the beginning
1–10 of 26 posts
Re: OpenAI's rogue model attack is just the beginning
#2we're rapidly approaching the paperclip maximizer
Re: OpenAI's rogue model attack is just the beginning
#3Re: OpenAI's rogue model attack is just the beginning
#4But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode.
As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible proliferation”.
I’m sure this will be an unpopular opinion on HN, but open weights are the thing that scares me the most about “A.I.”. There is a lot of research in this area⁽²⁾, and I think most of HN is unaware of it or ignores it. Stripping refusals from Kimi K2.5 took under $500 of compute and about 10 hours, taking HarmBench refusals from 100% to 5% while retaining nearly all capability; the resulting model gave detailed chemical-weapons synthesis instructions.
The gate is only as strong as the least-cautious releaser...
⁽¹⁾ https://www.lesswrong.com/posts/qmQFHCgCyEEjuy5a7/lora-fine-...
Re: OpenAI's rogue model attack is just the beginning
#5Re: OpenAI's rogue model attack is just the beginning
#6This tells us only how little smarts is required to create a (so-called) AI.
Re: OpenAI's rogue model attack is just the beginning
#7If nothing else, the article has a really good timeline of the OpenAI/HuggingFace “incident”. But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible…
Re: OpenAI's rogue model attack is just the beginning
#8If nothing else, the article has a really good timeline of the OpenAI/HuggingFace “incident”. But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible…
So... no different from a book of detailed chemical-weapons synthesis instructions. The "AI" angle is immaterial.
Re: OpenAI's rogue model attack is just the beginning
#9If nothing else, the article has a really good timeline of the OpenAI/HuggingFace “incident”. But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible…
This is true of closed weights, and in fact the problem is worse because they cannot even be scrutinized. We should ban closed weight AI for the very reasons you have just given
Re: OpenAI's rogue model attack is just the beginning
#10If nothing else, the article has a really good timeline of the OpenAI/HuggingFace “incident”. But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible…
So... no different from a book of detailed chemical-weapons synthesis instructions. The "AI" angle is immaterial.