For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. However, a model that always participates but "only" succeeds 10% of the time is golden -- just run it more often, or give it more tokens. Attackers are patient, and many of them are well-resourced.
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
21–30 of 50 posts
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#22The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…
This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#23I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_... these two claims can't be true at one and the same time: (a) they're distilling our secret sauce! (b) they'll never catch us!
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#24The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#25> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2]. This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 t…
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#26I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_... these two claims can't be true at one and the same time: (a) they're distilling our secret sauce! (b) they'll never catch us!
The fact the Opus 5 seems to be as capable as Fable/Mythos on everything except cyber, and Anthropic explicitly say they removed all offensive cyber training data, I think lends further credence to the idea that Mythos was designed from day zero to excel at offensive cyber capabilities.
If that’s true, then we would expect divergence in open models of their capabilities come from distillation, no frontier class cyber capable model has seen significant public availability. Which means there simply isn’t data to distill from.
It also calls into question the entire narrative around Mythos capabilities being a complete surprise for Anthropic, and an inevitable outcome of scaling up LLMs.
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#27The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#28Earlier quoted context omitted.
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#29Earlier quoted context omitted.
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#30Earlier quoted context omitted.
I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?
start here: https://huggingface.co/blog/mlabonne/abliteration