Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

421–430 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#425

Earlier quoted context omitted.

To make the discussion constructive, can you give specific reasons (ideally with examples) about why it is so useless for you? How exactly are you using it that you think any output from it can easily be replaced with a Wikipedia search?

The cybersecurity and bioweapons filters reach so far that they set in as soon as the model even glazes anything STEM-related. It might give a good impression of ones ex or write a decent fanfiction but anything that could bring humanity forward is strictly off-limits.

The filter is not simply a bioweapons filter: the model card seems to say that the filter triggers on anything related to biology or chemistry.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#426

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

There's no doubt in my mind they would if they could.

[dead]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#427

Earlier quoted context omitted.

I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml resear…

I think all this started with post opus 4.5, that's when claude started wrecking my shit without extreme oversight. Codebases it was making positive contributions to before were slowly and constantly being eroded and wrecked. Give it tasks in isolation? still does well, but the moment it sees the bigger picture, it goes to shit. I chalked it up to a bad model but this makes it all seem like it may have been by design…

Constraint decay is an issue with all LLM-based agentic development, at least for now.

Humans can maintain a long- and medium- term memory of constraints that they consciously (or subconsciously!) apply to the code that they write. The current crop of AIs are all amnesiacs, like the protagonist in Memento, falling back onto general instead of institutional knowledge.

For now, we are safe. We can rent out our meat brains for money for a little while longer.

Next year? Who knows...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#428
post #129

Earlier quoted context omitted.

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

To make an analogy: Imagine a patron gets banned from ordering alcohol at a particular establishment, because they got too drunk one time. It's completely reasonable for the establishment to reject a request for an alcoholic drink, and suggest something alcohol-free instead. It is not reasonable for them to say "sure, here's your alcoholic drink as you requested" and give them an alcohol-free substitute without telli…

> It is not reasonable for them to say "sure, here's your alcoholic drink as you requested" and give them an alcohol-free substitute without telling them.

Your analogy doesn't work because: - they tell you the rules at the entrance of the bar - they totally tell you when they give you a substitute

The only issue is the bartender asking you for your money before serving you the drink really but again, this is known since day 1 by the customers.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#429
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

The other major thing is almost as bad, and actually maybe even worse for trust of AI features in b2b apps:

> Anthropic requires 30 day data retention for Fable and Mythos

https://news.ycombinator.com/item?id=48464258

I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for training models."

That simple blanket statement is no longer true. Also, most normal people/customers only read headlines, and this is a huge story. From my point of view, as someone deploying LLMs in my apps, trust comms with my clients just got set back two years.

Post reply on HN