Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

461–470 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#462

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate?

From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#463

Earlier quoted context omitted.

We all need to use nuclear, bio and cybersec terms in all our code to make low quality filtering like this untenable. When you can't work on a resume that has cybersecurity or biology terms in it or reply to a job opening that includes them because the "AI" filtering is so bad that it confuses these for threats, that deserves a collective response, particularly to an IPO'ing company that claims they'll make workers o…

That's why I use M-x spook to generate all of my variable names

You can still find those clipper keyword storms in Usenet archives.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#464
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

The "tradeoff" warning implies they stand by their thinking and don't think there was anything qualitatively wrong with it which, if nothing else, is helpful so potential customers can know how they think. I think the core lesson is if you want reliable infrastructure to build into an application you should use a different provider. (edit: I'm not specifically an Anthropic hater, but having just spent some time addin…

Is it not a trade off? I think they made the wrong choice, but it seems reductive so say there was no choice at all and should never have been consideration of trade offs of silent versus not.

Even wide open, uncensored models are often the product of a deliberate choice. I have a hard time faulting people for intentionality (even when they get it wrong).

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#466

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?

They have no choice, enterprise customers won’t touch them unless they take a position like this. It’s a practical decision for them at the end of the the day.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#467
post #414

Earlier quoted context omitted.

Same with VISA/Mastercard deciding what we can/cannot buy. The only solution is to stop using their credit cards at all.

Yes, Monero is a lot better than credit cards for privacy and freedom. I hope to see it accepted more.

[dead]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#468
post #379

Popcorn for watching all those webapps being penetrated. Long live static websites without any Javascript.

this! javascript does add some nice UI&UX but i learned to do without, makes you get creative.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#470
So, this could have been implemented even before this Fable, could have been there from long ago. Puts a different perspective on all the reddit threads "opus is dumb today". Who knew that if you said the wrong word, the model would just intentionally feed you BS, without you even knowing it did.

WOW, never liked the virtue signaling Anthropic did with gov contracts but whatever. Got passed that. But this?

Post reply on HN