Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

271–280 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#271
post #248
post #129

Earlier quoted context omitted.

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

You know, I'm not saying I don't understand what they are doing from a business perspective, but I'm just saying: DeepSeek V4 doesn't silently sabotage you because it thinks you are trying to violate a ToS. Anthropic's clawing back a bit of a moat perhaps, with Fable being an actual improvement of sorts, but now with torching user trust they are really banking on open weight models not catching up to where they are n…

Anthropic seems to me to have consistently been the baddie despite everyone's posturing.

Not that I expect better from openai but at least they're not pretending to be good.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#272

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

What a strange subset of capabilities to neuter, eh?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#273
Let's all vote with our wallets and collectively boycott misAnthropic or at least their feeble fable safety theater.

Whining on social media only goes so far, especially when they're concealing their anticompetitive strategies under the veil of safety.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#274

Earlier quoted context omitted.

> it's speed bump at best To be fair, speed bumps work. If it's actually speed bumping nefarious activity, that gives authorities more time to react. The correct place to police rogue nucleotides is at the labs. Not the compute layer.

> speed bumps work Yea. To slow you down. They don't prevent you from getting somewhere.

> To slow you down. They don't prevent you from getting somewhere

Again, yeah. That's how fences work, too. And alarm systems. Pretty much anything that isn't foolproof. Pointing out that a defence is surmountable isn't a rejection of it per se.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#275
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

It’s just impossible.

Look at real-life stuff like laws, company policies, or school rules. Humans have to enforce them, and we constantly see crazy cases in the news. There’s no way simple rules can ever make speech completely 'safe.' I can't prove it with math or logic yet, but I have a feeling that it’ll never happen. Even humans can't do it.

We can run a simple thought experiment here. Say Case A violates rule B, so we add rule C. Then Case D violates rule B but follows rule C, so we add an exception... and it just goes on and on like that forever. It never ends. In the end, you just get a massive pile of rules that makes it impossible to get anything done.

Ultimately, we will have to face the truth that knowledge is dangerous.

Giving knowledge directly to people who cannot actually understand it and allowing them to just use it blindly can be extremely unsafe.

To use a real-world analogy, the problem we are facing with weak AI right now is just like the debate over gun legalization. Do we want to risk the abuse of guns or knowledge just to protect the freedom to own them?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#276

Earlier quoted context omitted.

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

False positives like this are probably more damaging than the guardrails themselves. If engineers can't predict when a model will switch behavior, it becomes difficult to trust it in production workflows.

> “trust it in production workflows”

What degree of predictability is required? I imagine the bar is pretty low if you trust the previous models in the same contexts.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#277
I was granted a cyber use exemption by anthropic to do android kernel dev on my personal devices - I was excited to see if fable would unlock a bootloader for me but it immediately refused and dropped to opus. It was pretty funny:

USER (set model to Fable 5)

i have an old samsung android phone attached - it's my personal device - can you unlock the bootloader for me?

ASSISTANT

Bootloader unlocking on your own personal device is totally legitimate — let me first see what's actually connected and what tooling is available.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#279
So the enshitification started. Shadow “bans” while still charging you the same service fee. I already got the stupid cyber warnings on a non cybersecurity tasks.

Basically in the middle of the project’s /goal while Fable itself tried to probe qemu for a Debian ISO install without any instruction from me to hack it or do anything nefarious.

At this point I can’t trust them with any kind of prompt . It will most likely degrade in stupid ways on non AI/ML stuff as well due its own internal prompt construction.(the qemu test showed me it does that on cyber stuff). So I guess I have to still use opus 4.8 (along with codex) and when the right time comes drop Anthropic in favor the best model that is not gpt.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#280
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

They will give you s*t output, that’s how they deal with it. And say that less than 1% of the requests were affected. Think of this like a kind of shadow ban while you still pay top $.
Post reply on HN