Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

131–140 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#131
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

I’m a noob about laws but isn’t this abusing its dominant market position and violates some antitrust law?

Why would it? There’s plenty of competition in the AI space.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#132
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

One year ahead of it's competition in what exactly? Vibe coding? From Opus 4.7 onwards each following model is becoming less useful as an assistant and turning you as the assistant. But I guess that's normal when it's trained to pass benchmarks end to end. In fact it has become extremely good at pushing against feedback with extremely convincing and intelligent takes, even when it's completely wrong. I have extensive…

Yeah, what's up with that. Lately I have found that it tries to find excuses to not do as told and instead do a totally different thing. I told it to write a yaml file according to some specifications and instead it coded a Python script to write the yaml...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#133

So I suspect Anthropic started A/B testing or just plain testing this a while ago, Tell HN: Claude flags biology / biotech questions https://news.ycombinator.com/item?id=47929885 Today, it's flagging population research questions, Using only the dataset you constructed, assess two questions: 1. **Mortality:** do [GROUP] show mortality that differs from (a) your comparison groups and (b) era- and sex-matched US popula…

I was digging into some orbital mechanics questions and I assume it decided I was trying to backyard-science my way into an orbital-bombardment weapon. Kind of wild how this product's impression has gone from "wow, this is pretty neat" to "irreverent sack of dog shit you" in 24 hours almost solely on the back of a half-baked moderation system.

Oh yes, also liquid propulsion systems. GNC stuff. All flagged.

I think LLMs are capable of intelligence amplification; and if you're in the subset of people who'd benefit from it the most, you'll get locked out.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#134
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

It royally pissed me off today by just continuing with credits without stopping to ask me if I was ok with it.

Ran up $30 in extra charges while it was just flashing on the screen that it was doing that after I walked away to do something while it was humming along.

It has always just told me I ran out of usage and had to wait before. Now? You’re just gonna pay extra because you left it unattended as you’ve done for the last year of use.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#135
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

> it won't just reject ML research, which I can understand I don't.

Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#136
Fable is a complete joke:

what's the best way to run this mcp server against the OData API used in this project? Can you come up with a PoC in a docker container?

https://github.com/oisee/odata_mcp_go

● I'll dig into two things in parallel: how this project talks to the OData API, and what the odata_mcp_go server needs to run. Let me start exploring.

Searched for 1 pattern (ctrl+o to expand)

● Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more ⎿ Tip: You can configure model switch behavior in /config

● Let me read the key integration files and fetch the MCP server's README at the same time.

  Read 2 files (ctrl+o to expand)
● Fetch(https://github.com/oisee/odata_mcp_go)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#137

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

It royally pissed me off today by just continuing with credits without stopping to ask me if I was ok with it. Ran up $30 in extra charges while it was just flashing on the screen that it was doing that after I walked away to do something while it was humming along. It has always just told me I ran out of usage and had to wait before. Now? You’re just gonna pay extra because you left it unattended as you’ve done for…

Do you have Usage credits turned on in your settings?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#138
At least Anthropic weren't lying when they said only a week ago or so "No one has figured out guardrails yet", because they apparently haven't either and Fable simply flat out rejects anything remotely connected to biology or security, no matter how trivial.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#139

Fable is a complete joke: what's the best way to run this mcp server against the OData API used in this project? Can you come up with a PoC in a docker container? https://github.com/oisee/odata_mcp_go ● I'll dig into two things in parallel: how this project talks to the OData API, and what the odata_mcp_go server needs to run. Let me start exploring. Searched for 1 pattern (ctrl+o to expand) ● Fable 5's safety measur…

And it charges you for that, and for when it decides to silently sabotage your request by routing to a dumbass model (without discount from Fable pricing)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#140
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

> It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition.

Making it look like you have something worth protecting is better for share prices than making something worth protecting.

Post reply on HN