Live data from Hacker News

Show HN: We post-trained a model that pen tests instead of refusing

argusred.com

31–40 of 49 posts

Re: Show HN: We post-trained a model that pen tests instead of refusing

#31

Relevant: https://news.ycombinator.com/item?id=48016224 what's the differnce between this vs running shannon on aws/bedrock fully airgapped in my vpc? I've got some pretty great results with shannon [no subprocessor and can pay via aws credits]. Even better using claude code token [effectively free with our $200/mo cc subscription] I tried kimi but it generally spins it's wheels extensively in it's thinking tokens. k…

It's named throughout our main website, the RL is on Kimi K2.6, benchmarks are vs K2.6: https://cosine.sh/blog/introducing-lumen-outpost. The ArgusRed page is a week old so it's not on there yet, but nothing's hidden. And K2.6 only needs attribution above a certain scale, the threshold Cursor hit and we haven't.

On Shannon airgapped in your VPC, if it works for you, you might not need us. A normal model will refuse or hedge on offensive tasks, we post-trained ours to just run the authorised stuff. For this one narrow job, a specialist that'll actually attack beats a generalist that won't.

Re: Show HN: We post-trained a model that pen tests instead of refusing

#32

> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.

Reminds me of a time when Tailscale cofounder went on a rant about how big bad AWS charges too much for bandwidth, and his solution was to send that money to Tailscale instead

Re: Show HN: We post-trained a model that pen tests instead of refusing

#33
post #30
post #27

IMO the most interesting thing about this is Kimi K2.6, an extremely capable model, can be relatively easily post-trained to allow pen tests. This in its own right proves that the defenses of Fable and others are temporary blocks, and AI based hacking is going to be effectively available to all parties regardless of stop gaps, as long as open models exist.

Agreed, and that's basically our premise. If a 5 person team can post-train an open model to do this, so can the people you don't want doing it, model-level refusals on open weights are a speed bump. Which is the argument for defenders having it too, not against.

literally anyone can "liberate" a foss model with access to weights

Re: Show HN: We post-trained a model that pen tests instead of refusing

#34
post #32

> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.

Reminds me of a time when Tailscale cofounder went on a rant about how big bad AWS charges too much for bandwidth, and his solution was to send that money to Tailscale instead

IIRC tailscale is directly P2P, sidestepping a large part of the infra costs...

Re: Show HN: We post-trained a model that pen tests instead of refusing

#35
post #32

Earlier quoted context omitted.

Reminds me of a time when Tailscale cofounder went on a rant about how big bad AWS charges too much for bandwidth, and his solution was to send that money to Tailscale instead

IIRC tailscale is directly P2P, sidestepping a large part of the infra costs...

And, as far as I'm aware, they don't charge for relay bandwidth even if you do end up needing it (which most users won't).

Re: Show HN: We post-trained a model that pen tests instead of refusing

#36
post #6

Earlier quoted context omitted.

I think the policy universally makes sense, who would want to give a tool like this to bad actors? But it does leave a big section of the market underserved. Particularly when Mythos was made accessible to very large orgs and then Fable was pulled on export grounds.

The problem is that it is a fool's errand to try to keep software tools from 'bad actors'. It is as pointless now as it was during the Crypto Wars. Information is simply too easy to move. https://en.wikipedia.org/wiki/Crypto_Wars

This is unrelated. The model is not being released directly - it's kept behind an API. You can't download the model and redistribute it like you can a piece of software, so the "information is simply too easy to move" ("information wants to be free") trope is a category error.

(don't mention distilling unless you understand why it's a different case than what's being described above)

Re: Show HN: We post-trained a model that pen tests instead of refusing

#37
post #32

> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.

Reminds me of a time when Tailscale cofounder went on a rant about how big bad AWS charges too much for bandwidth, and his solution was to send that money to Tailscale instead

That…isn’t the same thing at all, because your recount is factually incorrect. I’m not even a Tailscale user and I know that this isn’t equivalent.

Re: Show HN: We post-trained a model that pen tests instead of refusing

#38
post #6

Earlier quoted context omitted.

I think the policy universally makes sense, who would want to give a tool like this to bad actors? But it does leave a big section of the market underserved. Particularly when Mythos was made accessible to very large orgs and then Fable was pulled on export grounds.

The policy is repugnant. Whoever delivers the first frontier model as open weights to the world which lacks these moral guardrails will win. Stop thinking you know morals better than your users, or get out of the way so a competitor who respects your users more can serve them!

One doesn’t “get out of the way” for competitors, one is beaten by them. You just don’t know how to scroll past something you don’t like instead of going to the comments to complain about it.

Re: Show HN: We post-trained a model that pen tests instead of refusing

#40
post #6

Earlier quoted context omitted.

I think the policy universally makes sense, who would want to give a tool like this to bad actors? But it does leave a big section of the market underserved. Particularly when Mythos was made accessible to very large orgs and then Fable was pulled on export grounds.

A lot of bad actors are both technically sophisticated and have more than enough resources to post train their model. Morally I think it's still the right choice, but consequence wise I doubt it's going to make a big difference.

Bad actors tend to keep their internal tooling extremely private/proprietary.

As few/none would create a model as capable as anthropic/openai can - this choice to limit access does mean that most bad actors will be working with less capable models of varying quality.

While some will be able to fork DeepSeek and get comparable performance, it still reduces the number of bad actors with access to tools that would effectively accelerate their efforts.

So I suspect if you could measure the alternate universe timelines where everyone gets access to non-aligned foundation models vs. heavily restricted access, you’d probably find that in the near/medium terms the universe with restricted access probably sees less negative impact overall.

Long term it’ll be a wash either way (eventually Opus-level models will run on 20 watts) and hopefully Anthropic is correct in their predictions that LLMs will grant a strong defenders advantage in the long run.

Post reply on HN