Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
411–420 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#412Earlier quoted context omitted.
I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml resear…
> I would never know whether the models they launch will silently corrupt my output You never knew to begin with, now you have an explicit reason to realize this. Any black box run entirely out of your control, where you can never verify the output, is subject to the same suspicion.
“Fool me once, shame on you. Fool me twice, shame on me. Fool me three times, shame on both of us.” -- S. King
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#413Earlier quoted context omitted.
To make the discussion constructive, can you give specific reasons (ideally with examples) about why it is so useless for you? How exactly are you using it that you think any output from it can easily be replaced with a Wikipedia search?
The cybersecurity and bioweapons filters reach so far that they set in as soon as the model even glazes anything STEM-related. It might give a good impression of ones ex or write a decent fanfiction but anything that could bring humanity forward is strictly off-limits.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#414News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#415Earlier quoted context omitted.
If the guardrails were so useless, people wouldn't be complaining about them.
People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#416Earlier quoted context omitted.
> I would never know whether the models they launch will silently corrupt my output You never knew to begin with, now you have an explicit reason to realize this. Any black box run entirely out of your control, where you can never verify the output, is subject to the same suspicion.
True enough, but that is true for all the products I buy. I do not expect to control every product I own. For some I prefer to have more control, for others I just need something that works out of the box. There is always an initial bias for trust when you buy something otherwise you would not spend your hard earned money on it. “Fool me once, shame on you. Fool me twice, shame on me. Fool me three times, shame on bo…
Some things are more obscure than others. It's easier to trust and verify Office SaaS than AI SaaS. The determinism and obviousness of most other activities make them less susceptible to hidden interference. AI run by someone else is the next level of black box for users compared to most other objects or services we usually interact with.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#417Earlier quoted context omitted.
OpenAI has been the absolute worst about this, historically. I found myself having to change my queries because it refused to serve things it deemed insensitive.
Yes, that's true. Excluding Fable, OAI models are the most refusal heavy. However, I'd rather get a refusal than response with poisoned output. Since currently there's no way to verify if poisoning happened or not, I don't trust Anthropic anymore, regardless of what they say. But my trust towards OAI is also brittle - what if they also do it, or start doing it? I want to have a verifiable way to know that the prompt…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#418Earlier quoted context omitted.
OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.
Eh, I expect open Ai to follow suit. I suspect this is surprising to folk because they aren’t the ones busy figuring out how to use LLMs for illegal acts. In general, HN users focus on making stuff, and not the safety side of things, or the scale of harms being enabled via LLMs and generative AI. If you are on the safety side of things the ratio of misuse to fair use is inverted and everything is at scale. Transparen…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#419News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
They need to walk back a lot more. Unilaterally revoking zero-data retention, even for enterprise contracts that explicitly require that ? Nope. Fable is utterly unusable for any kind of security work. I tripped the safeguards yesterday - using Fable to dig into a complex (& annoying) security bug that has so far resisted both human and Opus 4.8 level investigation. "Sorry Dave, I can't let you do that." For the time…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#420News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
This is different to the cyber limitations though. To be precise - it makes the "won't work on frontier machine learning" refusal the same as the "won't work on cyber security" refusal (instead of the way it previously would work on frontier machine learning problems but give sub-optimal answers without informing the user)
Of course, it’s impossible to know if that was deliberate sabotage, or model misbehaviour. Which is exactly the problem.
That may be considered malware / a criminal act tbh.