Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

251–260 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#251
post #64

I make privacy tooling and Fable 5 rejects the vast majority of my prompts to analyze and improve the software that I've written. It's bleak.

Anthropic refused to let Fable analyze my own project's memory safety, the one thing I absolutely wanted it to do. Even Fable thought it was stupid.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#253

Earlier quoted context omitted.

Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.

Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?

More like concerned about distillation.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#254
post #123

Earlier quoted context omitted.

Anthropic is trying to hide bad behavior by being vague, it's important to not be vague when calling it out.

I'm of the opinion that removing guardrails is how you force regulation. What's your opinion on the balance?

They’re not safety guardrails they’re anthropic doesn’t like anyone who isn’t anthropic working on AI rails

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#255
post #197
post #171

Earlier quoted context omitted.

If it's a violation of ToS, just reject instead of silently downgrading.

But then someone would figure out some prompts that don't trigger this, and Anthropic wouldn't be able to try and disadvantage competitors.

Except they openly reject many many other classes of prompts, including extremely high stakes CBRN.

It's only the direction that has direct potential business impact they've decided to sabotage instead of reject.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#256
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

One thing is a model that's trained from the start to say "This topic is above my pay grade" to any mention of the status of Taiwan, etc.

Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics.

I bet a similar architecture is already deployed, e.g. to fight porn, planning of crimes, etc. But it can be turned into a dynamic system that provides controllable different answers (including unhelpful or misleading answers) based on geography, language, browser fingerprints, or the current political climate. All this could happen undetectedly and gradually if desired.

Welcome to a cyberpunk dystopia.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#257
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

I thought it was known since a few years now that if you train models to NOT do certain things, then they start behaving in weird ways…

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#258
The main thing that sucks with Claude is the extremely low limits before you get fail2banned for 6 hours. I'm out. Refund requested. Grok and Gemini Pro are way better with the throttling, can't comment on ChatGPT, haven't used that for a year.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#260
post #256
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

One thing is a model that's trained from the start to say "This topic is above my pay grade" to any mention of the status of Taiwan, etc. Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics. I bet a s…

This level of censorship kinda does make even Soviet or Maoist censors look like a honest straightforward bunch in comparison.

A very ironic result from a company supposedly valuing the opposite.

Post reply on HN