Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
541–550 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#542Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#543Earlier quoted context omitted.
The other major thing is almost as bad, and actually maybe even worse for trust of AI features in b2b apps: > Anthropic requires 30 day data retention for Fable and Mythos https://news.ycombinator.com/item?id=48464258 I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for tra…
> I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for training models." They claim they're not using it for training, only for "safety", and in fact I believe them. If you think they're lying, then why didn't you think they were lying about zero retention before? And "don'…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#544Earlier quoted context omitted.
Or they wanted the model to be good at these things, for the companies that legitimately need access to these capabilities.
so only the chosen for-profit companies by Anthropic are allowed to use frontier ai in the name of safety? what kind of joke is that? you people here can't be that dumb..
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#545Earlier quoted context omitted.
So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with... https://tylereaves.github.io/uk-rail-map/ This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling sche…
I love computational mapping projects, because there is this hard problem of which towns to show on the map. Your Scotland map shows towns without rail (although some had rail previously, like Callander, Aberfeldy), it prefers insignificant (population-wise) places while ignoring the larger cities next to it (Scone instead of Perth, Bannockburn instead of Stirling, Inverness is missing, Dundee is missing, Aberdeen is…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#546Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#547Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…
I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?
If someone is going to make a lot of money (fame, etc), it might as well be me.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#548Earlier quoted context omitted.
> I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for training models." They claim they're not using it for training, only for "safety", and in fact I believe them. If you think they're lying, then why didn't you think they were lying about zero retention before? And "don'…
I explained/ranted about why this new scenario is far more worrisome in this comment: https://news.ycombinator.com/item?id=48488781
Also, while I do agree that Anthropic's internal controls are unlikely to be on the level of AWS's or Azure's. I'm pretty confident they're good enough that random PMs aren't going to get access to things like that, especially for use in formal projects. Especially since "safety" is Anthropic's other obsession, which means "safety" data are going to be watched.
But anyway, we seem to be agreed that retaining stuff that used to get flushed early is a risk, and every copy is a risk, and sending it to more companies is a risk, regardless of the fine points of how things might go wrong.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#549Earlier quoted context omitted.
I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?
the problem is that the guardrails prevent us from performing real security work which is friction that is incurred by the legitimate user but not by a moderately sophisticated threat-actor. for example in my org it is part of the culture that security has no seat at the table. that is a separate problem, but the number of orgs like mine are more numerous than the number of orgs where security isn't a cost-center. we…
The right way to do it is to run a series of agents... many of which are nonetheless built on frontier models (and nearly none of which are built on some local 27B Qwen variant...). One thing the latest models are good at is orchestrating other agents.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#550Earlier quoted context omitted.
I explained/ranted about why this new scenario is far more worrisome in this comment: https://news.ycombinator.com/item?id=48488781
I still don't think "enterprise" customers' data have enough training value for an "over-eager PM" to bother with them. An over-eager PM obsessed with AGI, no less. I'm pretty sure more training on corporate slop is not the path to AGI (even less so if you want "aligned" AGI...). Also, while I do agree that Anthropic's internal controls are unlikely to be on the level of AWS's or Azure's. I'm pretty confident they're…
By over-eager PM, I didn't mean someone being malicious, just moving too fast to think about how some logging they set up might have negative effects for my clients way down the line.
Then months later, some other person finding a store of novel data, and being like... that looks nice... not gonna ask any questions/look a gift horse in the mouth... woohoo AGI!
On the far darker side, while I am a fan of the team at Anthropic: good intentions and all, they had to pay a $1.5B settlement for knowingly ingesting copyrighted books. That was just the cost of doing business. They did that, and now they are a trillion dollar company.