Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

201–210 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#201

Earlier quoted context omitted.

Some people made grandiose claims about its capabilities and I wanted to experience it myself.

OK, but for almost 24h straight? That seems a little obsessive, and not in the good way.

Getting excited about the announcement of new capabilities is very normal.

People used to wait in line all night to buy an iPhone. This isn’t that different.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#202

Earlier quoted context omitted.

> it won't just reject ML research, which I can understand I don't.

Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.

[deleted]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#203
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

[deleted]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#204
post #121

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

> We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose Whatever problem we might have with them, they explicitly say that they do not do this in the launch post.

[dead]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#205
post #165

Earlier quoted context omitted.

The challenge is the examples they’ve mentioned (distributed training infra? ML acceleration techniques?) go beyond what’s prohibited by their ToS and is like a catch net. I would wager the majority of ML and data science work in the world aren’t frontier LLM development.

Yes, this is the problem. They are business interests of Anthropic and have nothing to do with “safety”

Safety of their IPO

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#206
post #15

Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.

Yes, the miasma worm does this since the new Hades campaign.

Note that the 3rd wave now also uses a pth file in pypi packages that _search system wide_ for any index.js or .github/setup.js to find its own payload. It literally splits up the payload on purpose to avoid detection.

Mitigation Tool: https://github.com/cookiengineer/antimiasma

Technical Blog Post: https://cookie.engineer/weblog/articles/malware-insights-mia...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#207
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

I suspect it's even more expensive to run than they are charging for. These safeguards are just an excuse to get people to use it less, because it's not actually sustainable to use. They want to tempt people to consider them the leader, and it may actually be somewhat stronger, but too expensive to actually use at scale, so they nerf it by downgrading you constantly.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#208

So I suspect Anthropic started A/B testing or just plain testing this a while ago, Tell HN: Claude flags biology / biotech questions https://news.ycombinator.com/item?id=47929885 Today, it's flagging population research questions, Using only the dataset you constructed, assess two questions: 1. **Mortality:** do [GROUP] show mortality that differs from (a) your comparison groups and (b) era- and sex-matched US popula…

Ah it just flagged my water solubility question!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#209
post #161

Earlier quoted context omitted.

One year ahead of it's competition in what exactly? Vibe coding? From Opus 4.7 onwards each following model is becoming less useful as an assistant and turning you as the assistant. But I guess that's normal when it's trained to pass benchmarks end to end. In fact it has become extremely good at pushing against feedback with extremely convincing and intelligent takes, even when it's completely wrong. I have extensive…

They def not 1 year ahead, at most 2 weeks ahead until Openai releases theirs. This guy def a Anthropic shill and probably doesn't use any other LLMs.

I only said one year because I was thinking anthropic fans might downvote my post, I think they have a few months lead and are deluding themselves that they can get regulation to halt development and stay on top

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#210
I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for...

At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is better than saying "we can't make it perfect, so whoops, we aren't going to have any guardrails at all"? Especially because Anthropic did pretty extensive red-teaming of Mythos & Fable...

Post reply on HN