Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

501–510 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#501

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

Mythos is supposedly good at security research.

Local Qwen 3.6 27B can hardly debug 5 lines of CSS or copy a short snippet from A to B without mangling it.

It's not like you can use the local model for security research or engineering biological weapons.

If you have $200k maybe you can get the hardware to run the larger open source models, but even they are behind latest proprietary models.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#502
post #10

I’m a dumb question asker and I’m not happy about the guardrails. Would you believe I’ve asked 20 questions and haven’t talked to fable yet? Every single thing gets rerouted to 4.8.

some static words in AGENTS.md trigger it as well as some mcp servers.

Even using incognito on the web page keeps refusing.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#503

Earlier quoted context omitted.

guess who is pushing for those anti-privacy laws? hint: they're publicly traded

I have encountered enough such people to know that the really heavy push is coming from the police and secret service circles. These are the workplaces that attract all the wannabe Stasi types.

I am 100% convinced the reason laptops came with webcams as standard so early on, even when webcams were an expensive option, was because law enforcement needed to spy on people.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#504

Earlier quoted context omitted.

Ah, dang it. My college professors warned me about this: the Wikipedia page I read the other day is wrong!

Did you read a Wikipedia page, or did you read a LLM-generated summary? When I looked this number up yesterday the LLM summary claimed it was millions, but I opened the Anthropic post I was looking for and verified it was indeed just 150,000. Are you sure you weren't just being lazy and trusting the summary?

I said what I meant:

https://en.wikipedia.org/wiki/DeepSeek

> In February 2026, Anthropic accused DeepSeek of using thousands of fraudulent accounts to generate millions of conversations with Claude to train its own large language models.[57]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#505

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

Mythos is supposedly good at security research. Local Qwen 3.6 27B can hardly debug 5 lines of CSS or copy a short snippet from A to B without mangling it. It's not like you can use the local model for security research or engineering biological weapons. If you have $200k maybe you can get the hardware to run the larger open source models, but even they are behind latest proprietary models.

I asked local qwen 3.6 what language my project was written in. It was a Java project, and it came back with C#. So I guess its pretty close.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#506
post #414

Earlier quoted context omitted.

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

Same with VISA/Mastercard deciding what we can/cannot buy. The only solution is to stop using their credit cards at all.

Easy to say, but every bank I've had the (dis)pleasure of doing business with only ever issued a Visa or Mastercard so it's not really feasible to just "stop using them"

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#507
post #195

I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!

That's a slime mold, not a fungus A slime mold is actually a giant amoeba, entirely distinct from a fungus.

> That's a slime mold, not a fungus

Now you sound like Pl@ntNet identify: "This is not a plant! Maybe fungi?"

(Edit: It doesn't seem catch amoebae in the same way. It suggested Goldmoss instead, with 1% confidence.)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#508

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

The guard rails aren’t about blocking professional malware authors. It’s about enabling a significantly larger population that isn’t as talented in acquiring those capabilities. Very different threat model and just because it’s not effective in one area doesn’t mean there isn’t value in making it more difficult for random Joe Schmoe in building an atomic bomb even if a kid before had done so successfully and turned his garage into a radiation danger site

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#509

Earlier quoted context omitted.

I was reverse engineering a medical device, and had to do a lot of trickery to get Opus 4.5 - not even Fable/Mythos, Opus - not to trip up its fucking CBRN filter. What happened with Fable is basically what I feared when they announced those restrictions. They took the shitty Opus CBRN filter and made it even worse. I pity the fools trying to use Anthropic AIs for anything biotech.

Opus has been fine on proteomics and bioinformatics for me. I have never seen a Claude model refuse on such grounds before in the past. Claude is still the best IMO, but it feels like its most frustrating and grating aspects are not down to the model’s abilities, but the increasingly heavy hand of Anthropic expressing itself within the model . Fable’s comically useless responses almost seem like a cynical marketing t…

My personal suspicion is that it went "medical hardware -> high-throughput screening -> biorisk" in that old Opus case.

I like Anthropic's work, and I would be the first to argue against all the usual "it's all PR" whine. But there is a limit. And whoever made those fucking filters needs to be fired out of a cannon into the sun.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#510

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

The guard rails aren’t about blocking professional malware authors. It’s about enabling a significantly larger population that isn’t as talented in acquiring those capabilities. Very different threat model and just because it’s not effective in one area doesn’t mean there isn’t value in making it more difficult for random Joe Schmoe in building an atomic bomb even if a kid before had done so successfully and turned h…

In other words security by obscurity.
Post reply on HN