> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.
I think the policy universally makes sense, who would want to give a tool like this to bad actors? But it does leave a big section of the market underserved. Particularly when Mythos was made accessible to very large orgs and then Fable was pulled on export grounds.
Show HN: We post-trained a model that pen tests instead of refusing
11–20 of 49 posts
Re: Show HN: We post-trained a model that pen tests instead of refusing
#12Re: Show HN: We post-trained a model that pen tests instead of refusing
#13Why create an offensive tool rather than a repo-scanning tool? I can't think of any way to safely release an offensive tool publicly.
I am able to get Opus and Sonnet to function as a red team agent. We don’t have some crazy special sauce, just a lot of trial and error. Basically add enough context proving we own the code and running services that it will run attempts to compromise our services.
It found tons of stuff that was not found with just scanning the code. It found serious security issues that had been in productions for years that humans never found. They weren’t things that were accessible externally but serious enough that we are thrilled to have these tools.
I can say that Fable did refuse to function with our harness. I am worried that soon you have to be in the special club to do this stuff with the SOTA models. A small company like ours doesn’t get accepted to their programs that remove guardrails. Even though our CEO has found and disclosed vulnerabilities to multiple companies and holds a patent around federated authentication.
Re: Show HN: We post-trained a model that pen tests instead of refusing
#14Re: Show HN: We post-trained a model that pen tests instead of refusing
#15Fantastic. Could you share more details what it was like post-training a model?
Re: Show HN: We post-trained a model that pen tests instead of refusing
#16> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.
I think the policy universally makes sense, who would want to give a tool like this to bad actors? But it does leave a big section of the market underserved. Particularly when Mythos was made accessible to very large orgs and then Fable was pulled on export grounds.
Stop thinking you know morals better than your users, or get out of the way so a competitor who respects your users more can serve them!
Re: Show HN: We post-trained a model that pen tests instead of refusing
#17Show HN: We told Claude to generate a marketing page for a theoretical pentesting model
Re: Show HN: We post-trained a model that pen tests instead of refusing
#18Why create an offensive tool rather than a repo-scanning tool? I can't think of any way to safely release an offensive tool publicly.
Re: Show HN: We post-trained a model that pen tests instead of refusing
#19> This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises. So this is the same policy that Anthropic and OpenAI have, it is just based on your criteria rather than theirs.
To me it looks like copycat marketing more than a strongly held stance
Artificial scarcity, membership club criteria to make members feel special
Perhaps there is an organization that awards this “responsibility” behavior, the EU comes to mind but not lucrative enough
As far as engagement farming goes, it got us to engage and boost its reach, for something we might otherwise ignore with more benign language
Once I get the answers I will execute
Re: Show HN: We post-trained a model that pen tests instead of refusing
#20Show HN: We told Claude to generate a marketing page for a theoretical pentesting model
The tool is live, you can test it.
It’s just more “We’re so smart we invented the boogeyman, trust us” slop marketing that’s been happening since gpt-2