Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

401–410 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#402
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

I hope this has some answers [1]. It’s on the front page right now, but your frustration clearly seems to have some implicit answers that [1] is trying to answer.

[1] https://news.ycombinator.com/item?id=48477135

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#403
post #185

The question is: If biological, computer security, and ML research are so bad, why do they even train on the relevant data? The only answer that makes sense is they wanted the model to be competent and usable in these fields, just not by you , which is why they had to bolt on a badly functioning crippling device after the fact.

Remove the relevant data, and just enough of the data around it will remain that the AI will be able to close the gap if given relevant documentation.

Not to mention that those capabilities are inherently dual use. If you know how to write C safely, you know how to spot unsafe C.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#404
post #224
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

Hey guys, check out this technique https://github.com/0xSufi/fable-jailbreak/ It works with security audits and other workflows that are currently blocked.

I don't want my ANT account banned, going to try this on some Chinese "proxies".

But this also looks quite useful to understand how CC dynamic workflows work. Was thinking of implementing something similar in my homemade orchestration system.

Did you get claude itself to RE the dynamic workflows?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#405
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

They are still downgrading. They just aren't doing it silently. I don't know how big of a win that is? They still trained on everyone else's data without license or attribution but want to prevent someone else from doing the same thing to them. Some pretty audacious hypocrisy from Anthropic this week.

It is much more reasonable to do it in a visible / flagged way. At least you have visibility over the quality of service you get as a customer.

Silent treatment is a breach of trust, what you buy changes depending on the context based on the goals of the producer. It is like your computer silently blocking ads from competitors at the hardware level, which is crazy. I think they erred on the wrong side of things due to IPO pressure.

At least there is competition from multiple companies. Still it is best to have personal benchmarks for the domain you are working on to have a real evaluation of the value you get for the money/time you spent on these products. Without trust, that might be the only way forward to keep the companies honest.

This happens eventually in all sectors, a good magazine/website that does independent product evaluation is priceless. Sadly, the new ad-driven internet decimated those that worked great in the 90/00s. Still there are independent blogs that does some evaluation and that is better than nothing.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#406
post #98

Earlier quoted context omitted.

Their gap over Chinese models like GLM-5.1 is nowhere near 18 months. In many areas, it’s less than 6 months. The best closed models 18 months ago were worse than Qwen3.6.

These coding agent models only started getting useful in January. Before that they were difficult to control autocomplete, and not very smart. January was an inflection point, and no open weights model has crossed over that same threshold. This is definitely recursive self improvement territory, except that we're prohibited from participating. It feels like the capability gap is wider than before.

Have you tried deepseek V4? It costs pennies and is as good as Opus 4.6 (I found 4.7 to be a downgrade, and cancelled my claude subscription before 4.8).

The threshold has definitely been crossed.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#407
post #359

Earlier quoted context omitted.

Ironically making a stink about it online is likely to have a larger impact then using their dedicated feedback or support channels (which go to claude, not a person)

the feedback is for something mindless though, "we don't care about societal harms". I wonder the overlap between these commenters and tech maga people, eg crypto bros & Elon stans.

In this case, no overlap between me and tech maga / crypto bro / elon stan.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#408

Earlier quoted context omitted.

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml resear…

> I would never know whether the models they launch will silently corrupt my output

You never knew to begin with, now you have an explicit reason to realize this. Any black box run entirely out of your control, where you can never verify the output, is subject to the same suspicion.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#410

I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…

Guardrails against what? Rehashing public wikipedia information?

Execution matters, and they did a trurly horrible job that crippled their product to the point of being useless and a joke. Huge mistakes were made and im sure they regret it already, heads will roll.

Post reply on HN