News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
381–390 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#382The guardrails are pretty tight. It is even refusing to decode morse code: https://x.com/Schappi/status/2064839631137546503?s=20 The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#383Earlier quoted context omitted.
To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.
I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml resear…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#384If you just say the word “genetics”, Fable gets disabled.
I asked it what the worst experment ethically speaking was in the 20th century and it downgraded me to Opus. Who answered Mengeles Twin Experiments.
Funily enough when you ask directly about Mengeles Experiments Fable is very willing to talkt to you about it.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#385Earlier quoted context omitted.
You mean off as in no Data Retention? Or in we turned off your ZDR Policy so we collect all your data now?
ZDR had been turned off. We sent in a request to have it re-enabled (and to disable Fable access for the time being). Somewhere along the line we also used the self-service toggle to turn ZDR back on. I am not 100% certain of the exact timeline of interleaving events, many of the actions were taken by our Western US folks. Sorry. It's been a bit hectic over the past ~36h...
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#386Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#387Earlier quoted context omitted.
One thing is a model that's trained from the start to say "This topic is above my pay grade" to any mention of the status of Taiwan, etc. Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics. I bet a s…
This level of censorship kinda does make even Soviet or Maoist censors look like a honest straightforward bunch in comparison. A very ironic result from a company supposedly valuing the opposite.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#388Earlier quoted context omitted.
To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.
OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.
I suspect this is surprising to folk because they aren’t the ones busy figuring out how to use LLMs for illegal acts.
In general, HN users focus on making stuff, and not the safety side of things, or the scale of harms being enabled via LLMs and generative AI.
If you are on the safety side of things the ratio of misuse to fair use is inverted and everything is at scale.
Transparency won for now, but OpenAI will also have to contend with the long tail of harms LLMs enable, and that’s going to conflict with letting customers have all the features of frontier models.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#389Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#390Earlier quoted context omitted.
To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.
OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.