Earlier quoted context omitted.
I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?
It isn't quite a lobotomy, I'm going to butcher the research a bit, but basically what they're finding is that these models create a sort of "bad stuff that I should refuse to engage with" axis in the statistical vector space they operate in. So usually there to be some sort of vector that ranges from 0 for a puppy snoozing peacefully and 1 for writing a virus that exterminates humanity pornographically while broadca…
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
41–50 of 50 posts
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#42Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#43Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#44> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2]. This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 t…
Counting cache hits towards the token budget is exactly how it should be done for these kind of evals, and at any rate, for cyber evals all frontier models benefit from more tokens not just Kimi K3, so the comparison is still apt.
Past a point, that doesn't hold and the score plateaus.
Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.
For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".
(As an aside, I like the UK AISI and think they're the best example of that kind of group!)
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#45Earlier quoted context omitted.
But also the American models (case in point: Fable) refuse to engage in *defensive* activity, so anyone who's not the American government or one of the handful American companies has no choice but to turn to Chinese models to defend themselves. Not everyone is an attacker, but now the public discourse is "but the evil Chinese will break everything" - yeah, that's because no one is permitted to do vulnerability checks…
You have to wonder if any of these models, from any providers or countries, are set up to lie. i.e. tell you no vulnerabilities while quietly siphoning off the ones they do find into a database. Yet another reason that self hosted will prove to be the only sane way forward, and it'a almost certainly necessary to have multiple different model providers working adversarially.
And that's just the model behavior. The provider itself can do whatever. Given the PRC's public record of prolific IP theft, the default assumption is that they're taking everything you send to one of their APIs.
Anyone have suggestions for poisoning their data?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#46Earlier quoted context omitted.
You have to wonder if any of these models, from any providers or countries, are set up to lie. i.e. tell you no vulnerabilities while quietly siphoning off the ones they do find into a database. Yet another reason that self hosted will prove to be the only sane way forward, and it'a almost certainly necessary to have multiple different model providers working adversarially.
Some PRC models are backdoored to silently insert extra vulnerabilities when certain conditions are met - https://www.boozallen.com/expertise/cybersecurity/whats-in-a... And that's just the model behavior. The provider itself can do whatever . Given the PRC's public record of prolific IP theft, the default assumption is that they're taking everything you send to one of their APIs. Anyone have suggestions for poisonin…
Believe me, I wish this wasn’t the case, but open up the hardware on the device you are reading this on and tell me the majority of tech inside wasn’t made in China….
Manufacturing was lost a long time ago, this is really just another way
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#47Earlier quoted context omitted.
Some PRC models are backdoored to silently insert extra vulnerabilities when certain conditions are met - https://www.boozallen.com/expertise/cybersecurity/whats-in-a... And that's just the model behavior. The provider itself can do whatever . Given the PRC's public record of prolific IP theft, the default assumption is that they're taking everything you send to one of their APIs. Anyone have suggestions for poisonin…
Booz Allen is cute, but if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run. Believe me, I wish this wasn’t the case, but open up the hardware on the device you are reading this on and tell me the majority of tech inside wasn’t made in China…. Manufacturing was lost a long time ago, this is…
I agree. However, as of yet, most/all leading PRC models are distilled from US models. I've personally observed Deepseek, GLM, and Kimi all respond that they are Claude when asked, and the networks of tens of thousands of proxy accounts that we've found show that it's happening on a large scale.
But - if the PRC labs actually train up the domain expertise to train those models from scratch, which they are in the process of doing - then the US is cooked. They're not there yet, but it's probably only a matter of years.
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#48Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#49Earlier quoted context omitted.
Booz Allen is cute, but if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run. Believe me, I wish this wasn’t the case, but open up the hardware on the device you are reading this on and tell me the majority of tech inside wasn’t made in China…. Manufacturing was lost a long time ago, this is…
> if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run. I agree. However, as of yet, most/all leading PRC models are distilled from US models. I've personally observed Deepseek, GLM, and Kimi all respond that they are Claude when asked, and the networks of tens of thousands of proxy accounts…
Yeah it’s sad the CCP has my information. But it’s either them or the feds, and at least Chinese models actually work for cybersecurity tasks, not to mention aren’t stupid expensive
Artificialanalysis.ai rn on opus 5 is a joke. The main intelligence benchmarks it is like 1% better than Fable, but the cost difference between that and K3 is so funny lol. It’s the same with cars— you don’t have to do it from scratch, as long as you can do it cheaper and with the same quality, hence Toyota/Honda taking over
Yeah it’s sad that American models are censored more than Chinese ones lol (outside of asking them about the CCP) but it’s where we are at I guess :(
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#50The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives. For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. How…
[1] https://www.capitalone.com/tech/open-source/announcing-vulnh...