Live data from Hacker News

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

nist.gov

11–20 of 50 posts

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#11
post #9

Earlier quoted context omitted.

NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and i…

What is the rationale for gpt-5.5 when gpt-5.6-sol exists?

[deleted]

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#13
post #6

It hints that the original Mythos was tuned/trained for cyber attacks.

Most large security companies were collaborating with Anthropic as well as OpenAI on developing what has become Mythos and GPT-5 for a couple years now. Also, most HNers have never actually played around with the unrestricted models - once you get past the initial hump of re-tuning harnesses it can be fairly powerful. HN never really had a prominent security userbase at the best of times, and it's gotten worse since.…

> most HNers have never actually played around with the unrestricted models

Because they can't, of course, so regardless of how accurate the benchmarks or claims about the model are, they're functionally irrelevant to most of us.

"Thing you don't have or that randomly restricts you is actually better than thing you do have and can use" may be true and is still a practically worthless claim for anyone wanting to get real work done.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#15

Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published. Am I reading this wrong?

The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#16
post #10

> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2]. This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 t…

The gap seems to be so small I'm not sure how much we should care. If we assume that the trend in the graph does continue as a rough linear improvement, the open models will have reached the same level as the present closed models in 6 months and likely saturated the benchmark by mid next year. That seems more important than where we are now and the exact level of measurement accuracy in July 2026. There is a difference between US and Chinese models but it doesn't look like it is going to be strategically significant.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#17

This again shows the differences in breadth of capabilities that scale offers. On public benchmarks for "regular tasks" the chasing models come close, but on closed ones they lag behind. Just looking at the Elo differences, k3 is at ~2000 Elo, compared to SotA closed models at 3000 Elo. That is a huge difference. Also, even on the public benchmarks, k3 only scores in the "low hanging fruit tasks", with 0 successful c…

Isn’t it just because they hit the token limit?

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#18
post #15

Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published. Am I reading this wrong?

The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...

Ah, miss that one. But it still looks incomplete:

"Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations."

... and

"Kimi K3’s overall cyber capability [...] was estimated from a single benchmark (ExploitBench"

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#19
post #9

Earlier quoted context omitted.

NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and i…

What is the rationale for gpt-5.5 when gpt-5.6-sol exists?

I find 5.6-sol quite good for coding, but for legal work it has a bias favouring big corporations - it tends to weaken your arguments and make document legally unsafe. 5.5 was much better.

Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

#20
I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_...

these two claims can't be true at one and the same time:

(a) they're distilling our secret sauce!

(b) they'll never catch us!

Post reply on HN