Earlier quoted context omitted.
NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and i…
What is the rationale for gpt-5.5 when gpt-5.6-sol exists?
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
11–20 of 50 posts
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#12Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#13It hints that the original Mythos was tuned/trained for cyber attacks.
Most large security companies were collaborating with Anthropic as well as OpenAI on developing what has become Mythos and GPT-5 for a couple years now. Also, most HNers have never actually played around with the unrestricted models - once you get past the initial hump of re-tuning harnesses it can be fairly powerful. HN never really had a prominent security userbase at the best of times, and it's gotten worse since.…
Because they can't, of course, so regardless of how accurate the benchmarks or claims about the model are, they're functionally irrelevant to most of us.
"Thing you don't have or that randomly restricts you is actually better than thing you do have and can use" may be true and is still a practically worthless claim for anyone wanting to get real work done.
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#14Am I reading this wrong?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#15Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published. Am I reading this wrong?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#16> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2]. This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 t…
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#17This again shows the differences in breadth of capabilities that scale offers. On public benchmarks for "regular tasks" the chasing models come close, but on closed ones they lag behind. Just looking at the Elo differences, k3 is at ~2000 Elo, compared to SotA closed models at 3000 Elo. That is a huge difference. Also, even on the public benchmarks, k3 only scores in the "low hanging fruit tasks", with 0 successful c…
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#18Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published. Am I reading this wrong?
The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...
"Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations."
... and
"Kimi K3’s overall cyber capability [...] was estimated from a single benchmark (ExploitBench"
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#19Earlier quoted context omitted.
NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI... Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**]. [**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and i…
What is the rationale for gpt-5.5 when gpt-5.6-sol exists?
Re: UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
#20these two claims can't be true at one and the same time:
(a) they're distilling our secret sauce!
(b) they'll never catch us!