Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

231–240 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#232

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

"the tok/s is unmatched for a wiki article pull." This is absolutely wonderful, thank you for making my day!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#233
I asked a question about an openssl s_client parameter and warned me that I need to talk to Opus about cybersecurity lol. FWIW I dont see much improvement and still see quite the same old annoyances, so far I would not pay extra for this for my usage.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#234
post #9

Earlier quoted context omitted.

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

Specifically only ML research

Aah my mistake. I had missed that ML had separate trigger behavior from cybersecurity/etc... Thanks.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#235

I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…

The "guardrails" are just Anthropic's attempt at building a moat. Guarantee they'll be seeking regulation around AI as well to ensure a form of regulatory capture. Guardrails, in this context, are useless. Anyone who's sufficiently motivated will either get around them, or will just run their own model on their home hardware. There's already tools that one can use to remove the guardrails present in open weight models.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#236
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

Or if GPU companies detected you were trying to train a model and injected intentional numerical errors.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#237
post #176

Earlier quoted context omitted.

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

Consumer GPS is still disabled at high speeds. I would argue the analogy doesn't carry due to harm and error rate differences.

Yep a totally different use case and set of guardrails. There’s very little (not zero) consumer utility in GPS above say 15k feet AND 400 MPH or whatever the actual limit is. That’s basically tracking model rockets that are incidentally impacted and nothing else, from what I can think of.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#238

Earlier quoted context omitted.

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

> used to When’d that change?

He’s probably thinking of the accuracy limit to civilians it launched with.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#239
post #95
post #64

I make privacy tooling and Fable 5 rejects the vast majority of my prompts to analyze and improve the software that I've written. It's bleak.

Why is this surprising or a problem?! It's a model demo, & their reasoning is reasonable and fair. Why all this drama.

Tech demo + theres the ability to provide feedback right at the answer interface if using the UI.

Provide feedback in the negative, a brief explanation, and move on with your day. It will improve with feedback, not with whinging into the void.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#240
Software engineers shouldnt be happy either. If model silently sabotage cybersecurity research of others software there is abdolutely no way to be sure it wont be sabotaging cybersecurity of AI slop code it generated yesterday.

This is bad precedent and no one wants to pay X to generate code to then have to pay X*10 to figure out why your company just got hacked.

Post reply on HN