I make privacy tooling and Fable 5 rejects the vast majority of my prompts to analyze and improve the software that I've written. It's bleak.
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
251–260 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#252Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#253Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#254Earlier quoted context omitted.
Anthropic is trying to hide bad behavior by being vague, it's important to not be vague when calling it out.
I'm of the opinion that removing guardrails is how you force regulation. What's your opinion on the balance?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#255Earlier quoted context omitted.
If it's a violation of ToS, just reject instead of silently downgrading.
But then someone would figure out some prompts that don't trigger this, and Anthropic wouldn't be able to try and disadvantage competitors.
It's only the direction that has direct potential business impact they've decided to sabotage instead of reject.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#256The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics.
I bet a similar architecture is already deployed, e.g. to fight porn, planning of crimes, etc. But it can be turned into a dynamic system that provides controllable different answers (including unhelpful or misleading answers) based on geography, language, browser fingerprints, or the current political climate. All this could happen undetectedly and gradually if desired.
Welcome to a cyberpunk dystopia.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#257Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#258Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#259Chat paused. Fable 5's safety features have flagged this chat.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#260The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
One thing is a model that's trained from the start to say "This topic is above my pay grade" to any mention of the status of Taiwan, etc. Quite another is an architecture where the big model is not mutilated, but is gaslighted. A different, simpler model checks the incoming prompt and alters it if it contains banned topics. Another simpler model checks the output and censors it if it contains banned topics. I bet a s…
A very ironic result from a company supposedly valuing the opposite.