Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

381–390 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#381
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again, so, once the bad behavior starts, you leave. Their apologies are as hollow as their moral posturing.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#382

The guardrails are pretty tight. It is even refusing to decode morse code: https://x.com/Schappi/status/2064839631137546503?s=20 The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...

Lol i can't even ask this sonnnet it imediately shuts down. What a ajoke

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#383

Earlier quoted context omitted.

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml resear…

I think all this started with post opus 4.5, that's when claude started wrecking my shit without extreme oversight. Codebases it was making positive contributions to before were slowly and constantly being eroded and wrecked. Give it tasks in isolation? still does well, but the moment it sees the bigger picture, it goes to shit. I chalked it up to a bad model but this makes it all seem like it may have been by design in retrospect.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#384

If you just say the word “genetics”, Fable gets disabled.

Yeah just tried it can confirm thats absolutely hilarious.

I asked it what the worst experment ethically speaking was in the 20th century and it downgraded me to Opus. Who answered Mengeles Twin Experiments.

Funily enough when you ask directly about Mengeles Experiments Fable is very willing to talkt to you about it.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#385
post #372

Earlier quoted context omitted.

You mean off as in no Data Retention? Or in we turned off your ZDR Policy so we collect all your data now?

ZDR had been turned off. We sent in a request to have it re-enabled (and to disable Fable access for the time being). Somewhere along the line we also used the self-service toggle to turn ZDR back on. I am not 100% certain of the exact timeline of interleaving events, many of the actions were taken by our Western US folks. Sorry. It's been a bit hectic over the past ~36h...

JFC, thats a terrible situation. Thats literally a lawsuit or multiple waiting to happen. Godspeed you seem to have had a few interesting days so far.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#387
post #256

Earlier quoted context omitted.

One thing is a model that's trained from the start to say "This topic is above my pay grade" to any mention of the status of Taiwan, etc. Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics. I bet a s…

This level of censorship kinda does make even Soviet or Maoist censors look like a honest straightforward bunch in comparison. A very ironic result from a company supposedly valuing the opposite.

I would claim the difference between being rejected an API request and being potentially jailed/shot is significant.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#388
post #360

Earlier quoted context omitted.

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.

Eh, I expect open Ai to follow suit.

I suspect this is surprising to folk because they aren’t the ones busy figuring out how to use LLMs for illegal acts.

In general, HN users focus on making stuff, and not the safety side of things, or the scale of harms being enabled via LLMs and generative AI.

If you are on the safety side of things the ratio of misuse to fair use is inverted and everything is at scale.

Transparency won for now, but OpenAI will also have to contend with the long tail of harms LLMs enable, and that’s going to conflict with letting customers have all the features of frontier models.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#390
post #360

Earlier quoted context omitted.

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.

OpenAI has been the absolute worst about this, historically. I found myself having to change my queries because it refused to serve things it deemed insensitive.
Post reply on HN