Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

391–400 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#391
post #195

I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!

That's a slime mold, not a fungus A slime mold is actually a giant amoeba, entirely distinct from a fungus.

Careful with that dangerous knowledge, you’ll end up in a list.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#392

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

>I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. I dont understand. This is just hyperbole right? The outputs are basically infinite and wikipedia most certainly isnt infinite.

The decimals of 1/3 are infinite as well and they don't contain a better-than-wikipedia article.

And even if they did, it would be useless if it's buried in useless data and your chances or pulling it are effectively zero.

This is regardless of the general discussion, just pointing that your argument isn't solid.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#393

Earlier quoted context omitted.

> To slow you down. They don't prevent you from getting somewhere Again, yeah. That's how fences work, too. And alarm systems. Pretty much anything that isn't foolproof. Pointing out that a defence is surmountable isn't a rejection of it per se .

Fences and speed bumps are hilarious defences if we are supposed to believe AI companies about the dangers of this technology. Having no safeguards is probably safer than having safeguards which do nothing but create a false sense of security.

Idk, whether we believe them or not, I believe the life scientists who are calling for regulation around the labs that produce DNA sequences. If they’re concerned, regardless of whether I trust the AI labs, speed bumps could help by giving those scientists a reasonably window in which to be notified and act.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#394
Fable 5 reminds me of the time when Claude models where att version 1 and 2. They were fresh competitors to ChatGPT, for those who gave Claude a try experienced it to be almost unusable because of how heavily guardrailed it was.

This time, Fable 5 comes with another surprise, it can intentionally sabotage for you instead of rejecting the prompt. How is this possible for Anthropic to be able to treat their customers like this? It’s because you guys allowed it to. No matter what Anthropic does, you keep paying for their services. Vote with your wallet.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#395
post #360

Earlier quoted context omitted.

OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model. But if Anthropic gets their way with regulatory capture, this could be the only future we'll see. To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.

OpenAI has been the absolute worst about this, historically. I found myself having to change my queries because it refused to serve things it deemed insensitive.

Yes, that's true. Excluding Fable, OAI models are the most refusal heavy. However, I'd rather get a refusal than response with poisoned output.

Since currently there's no way to verify if poisoning happened or not, I don't trust Anthropic anymore, regardless of what they say.

But my trust towards OAI is also brittle - what if they also do it, or start doing it?

I want to have a verifiable way to know that the prompt I sent was the prompt the model received. I want to know if anything was injected as well - I understand they may not necessarily be able to reveal the exact steering, but at least give me the steering category and its hash or something.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#396

Earlier quoted context omitted.

One year ahead of it's competition in what exactly? Vibe coding? From Opus 4.7 onwards each following model is becoming less useful as an assistant and turning you as the assistant. But I guess that's normal when it's trained to pass benchmarks end to end. In fact it has become extremely good at pushing against feedback with extremely convincing and intelligent takes, even when it's completely wrong. I have extensive…

Yeah, what's up with that. Lately I have found that it tries to find excuses to not do as told and instead do a totally different thing. I told it to write a yaml file according to some specifications and instead it coded a Python script to write the yaml...

I got a worrying one: a day after getting opus 4.8, I tasked CC to add specific TXT records to our subdomain.example.com as per ticket I've received. CC has access to that ticket via Atlassian MCP, and started doing terraform code changes in a local git branch. Somewhere along the way it said that to do that it needs an approval from a company's VP (ticket requester) as "subdomain.example.com" is critical (it isn't). Then it refused to open a pull request, immediately deleted the local git branch along with all the changes and refused to proceed without evidence of approval from that VP. No amount of explaining, then pleading, and then threatening moved it. It was surreal and I was shocked and frankly pissed. It was amusing in the end because the day earlier it had no problem adding those same TXT records to example.com. Codex did those changes in 1/4 of time and no complaining.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#397

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…

Did my Claude get permanently dumber today because I asked fable to assess my Fairplay integration?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#398
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

The “1 year” part is key - all these safeguards etc are basically nonsense because in a few years at most one of the Chinese labs will release something equivalent, and in 10 years you’ll be able to run it locally with absolutely no safeguards at all

I think you're very optimistic with the "a few years", I'm confident all of the parties building AI models are working on Mythos equivalents / competitors, and if they can undercut Anthropic by making it more widely available and / or affordable they will. I give it three months tops. In a year all the major players will have an equivalent. In three years it'll be widely available, as more and more AI focused datacenters go online.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#399
post #185

The question is: If biological, computer security, and ML research are so bad, why do they even train on the relevant data? The only answer that makes sense is they wanted the model to be competent and usable in these fields, just not by you , which is why they had to bolt on a badly functioning crippling device after the fact.

Is what you suggest about training even possible? Most exploitation techniques are really just about having in-depth knowledge of how components work. For example, I imagine a sufficiently powerful model could fairly easily re-invent the ROP chain from first principles if it just knew how the stack works. This same principle applies to much more complex attack too; exploitation is often just an exercise in knowing va…

It would still degrade it's effectiveness, which is what they claim to want. Exaggeratedly: If it wasn't so, you'd just need fundamental math in the training data, as everything else can be derived.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#400

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

I work on software that talks to mass spectrometers and it consistently refuses to refactor even an input file parser, presumably because it can infer it’s related to biology? Useless indeed.

I was reverse engineering a medical device, and had to do a lot of trickery to get Opus 4.5 - not even Fable/Mythos, Opus - not to trip up its fucking CBRN filter.

What happened with Fable is basically what I feared when they announced those restrictions. They took the shitty Opus CBRN filter and made it even worse.

I pity the fools trying to use Anthropic AIs for anything biotech.

Post reply on HN