Earlier quoted context omitted.
A lot of bad actors are both technically sophisticated and have more than enough resources to post train their model. Morally I think it's still the right choice, but consequence wise I doubt it's going to make a big difference.
Bad actors tend to keep their internal tooling extremely private/proprietary. As few/none would create a model as capable as anthropic/openai can - this choice to limit access does mean that most bad actors will be working with less capable models of varying quality. While some will be able to fork DeepSeek and get comparable performance, it still reduces the number of bad actors with access to tools that would effec…
Show HN: We post-trained a model that pen tests instead of refusing
41–49 of 49 posts
Re: Show HN: We post-trained a model that pen tests instead of refusing
#42IMO the most interesting thing about this is Kimi K2.6, an extremely capable model, can be relatively easily post-trained to allow pen tests. This in its own right proves that the defenses of Fable and others are temporary blocks, and AI based hacking is going to be effectively available to all parties regardless of stop gaps, as long as open models exist.
Re: Show HN: We post-trained a model that pen tests instead of refusing
#43Earlier quoted context omitted.
A lot of bad actors are both technically sophisticated and have more than enough resources to post train their model. Morally I think it's still the right choice, but consequence wise I doubt it's going to make a big difference.
Bad actors tend to keep their internal tooling extremely private/proprietary. As few/none would create a model as capable as anthropic/openai can - this choice to limit access does mean that most bad actors will be working with less capable models of varying quality. While some will be able to fork DeepSeek and get comparable performance, it still reduces the number of bad actors with access to tools that would effec…
I wanted to create a harness with a collection of memories in order to play the upcoming downunderctf. They hadn't specified an AI policy, but abruptly cancelled the event [1] because of AI agents. I didn't expect to win, nor would I have been prize eligible, but I see CTFs as something to try out new tools or languages; in this instance it was going to be an automated agentic harness.
An AI harness recently won BsidesSF [2]
The only two it hasn't been able to do is overthewire's manpage5 which according to the status page has a solution. And drifter3 which I don't know if it currently has valid a solution. (Vortex13 and formulaone3 currently don't have valid solutions).
[0] https://en.wikipedia.org/wiki/Capture_the_flag_(cybersecurit...
[1] https://xcancel.com/DownUnderCTF/status/2062802249173356753#...
Re: Show HN: We post-trained a model that pen tests instead of refusing
#44Why create an offensive tool rather than a repo-scanning tool? I can't think of any way to safely release an offensive tool publicly.
You need both, scanning for your own code, pen testing to actually prove vulnerabilities, otherwise it can be very noisy and one of the things that most tools currently suffer from is they give you too many false positives. For the moment. The pen testing we gated it for now until we resolve the debate of safety.
I get that both need to exist as tools. I just don't see any safe way of doing a truly public release of the offensive end of it, you'd need to coordinate with established entities somehow.
Re: Show HN: We post-trained a model that pen tests instead of refusing
#45Re: Show HN: We post-trained a model that pen tests instead of refusing
#46>no benchmarks on standard Cyber benchs
ok