Live data from Hacker News

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

nist.gov

21–30 of 50 posts

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#21
The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives.

For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. However, a model that always participates but "only" succeeds 10% of the time is golden -- just run it more often, or give it more tokens. Attackers are patient, and many of them are well-resourced.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#22

The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…

> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity

This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#23

I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_... these two claims can't be true at one and the same time: (a) they're distilling our secret sauce! (b) they'll never catch us!

Statements of the form “X can never happen” tend to be weak, so that’s a bit of a strawman. But, in spirit, of course they can both be true. There are matters of degree and the intervention of countermeasures.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#24

The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…

> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.

I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#25
post #10

> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2]. This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 t…

Counting cache hits towards the token budget is exactly how it should be done for these kind of evals, and at any rate, for cyber evals all frontier models benefit from more tokens not just Kimi K3, so the comparison is still apt.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#26

I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_... these two claims can't be true at one and the same time: (a) they're distilling our secret sauce! (b) they'll never catch us!

I would read the divergence as mostly evidence that Mythos et al, had offensive cyber capabilities as part of their RL training. I.e. they’re models specifically trained to be good at offensive cyber, rather than being general purpose models that happened to become good at offensive cyber via emergent behaviour from sheer scale.

The fact the Opus 5 seems to be as capable as Fable/Mythos on everything except cyber, and Anthropic explicitly say they removed all offensive cyber training data, I think lends further credence to the idea that Mythos was designed from day zero to excel at offensive cyber capabilities.

If that’s true, then we would expect divergence in open models of their capabilities come from distillation, no frontier class cyber capable model has seen significant public availability. Which means there simply isn’t data to distill from.

It also calls into question the entire narrative around Mythos capabilities being a complete surprise for Anthropic, and an inevitable outcome of scaling up LLMs.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#27

The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…

[flagged]

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#28

Earlier quoted context omitted.

> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.

I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?

start here: https://huggingface.co/blog/mlabonne/abliteration

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#29

Earlier quoted context omitted.

> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.

I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?

You can find the current state-of-art tool for censorship removal here: https://github.com/p-e-w/heretic

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#30

Earlier quoted context omitted.

I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?

start here: https://huggingface.co/blog/mlabonne/abliteration

Note that this describes an older, currently pretty much obsolete technique which does lobotomize the model somewhat.
Post reply on HN