Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
461–470 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#462Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…
From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#463Earlier quoted context omitted.
We all need to use nuclear, bio and cybersec terms in all our code to make low quality filtering like this untenable. When you can't work on a resume that has cybersecurity or biology terms in it or reply to a job opening that includes them because the "AI" filtering is so bad that it confuses these for threats, that deserves a collective response, particularly to an IPO'ing company that claims they'll make workers o…
That's why I use M-x spook to generate all of my variable names
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#464News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
The "tradeoff" warning implies they stand by their thinking and don't think there was anything qualitatively wrong with it which, if nothing else, is helpful so potential customers can know how they think. I think the core lesson is if you want reliable infrastructure to build into an application you should use a different provider. (edit: I'm not specifically an Anthropic hater, but having just spent some time addin…
Even wide open, uncensored models are often the product of a deliberate choice. I have a hard time faulting people for intentionality (even when they get it wrong).
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#465Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#466Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…
I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#467Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#468Popcorn for watching all those webapps being penetrated. Long live static websites without any Javascript.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#469Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#470WOW, never liked the virtue signaling Anthropic did with gov contracts but whatever. Got passed that. But this?