Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

551–560 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#551

Earlier quoted context omitted.

I think it’s safe to say that many of us feel a lot less safe directly because of these policies and the inferred intentions of the company behind them. Nobody is arguing for unsafe models. We just don’t want to live in the plot of Deus Ex.

> Nobody is arguing for unsafe models Then what are people arguing for? I see only two totally distinct options: unsafe models or someone being the arbiter of safety

I’d settle for a different someone at this point. What were you expecting me to say?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#552
post #197
post #171

Earlier quoted context omitted.

If it's a violation of ToS, just reject instead of silently downgrading.

But then someone would figure out some prompts that don't trigger this, and Anthropic wouldn't be able to try and disadvantage competitors.

Ultimately Anthropic has to compete within the bounds of the law, even when doing an anti-competitive thing would make it easier for them to compete.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#553
post #340

Earlier quoted context omitted.

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

Didn't early GPS have fudge factor on the most precise bits? As such you could only get to a few meters of accuracy. Not critical for sea navigation or even to general positioning when paper maps were still used.

The term of art here is "Selective Availability" and the added error margin was up to 100 meters.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#554
post #453

Earlier quoted context omitted.

How would the AI be able to find the footage itself?

I use Codex and wanted it to sort through the footage and use subagents to review. Codex limits are fairly generous, esp paired with mini models for this kind of task generally, but even GPT5.5 usage is still pretty generous. Again, it’s the only refusal I’ve gotten for coding/agentic tasks, and it has a basis in law somewhere, so I don’t fault OpenAI for that.

Very cool, thanks.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#555

Earlier quoted context omitted.

Your rebuttle seems to be arguing it's okay for a bartender to simultaneously say: "This is alcohol" And "Or maybe it isn't alcohol." Or to rephrase it, "They tell you the rules at the entrance, they then tell you they don't follow those rules and they are totally serving alcohol even if they are not."

No they tell you at the entrance that at any point they may unilaterally decide to replace the alcoholic drink you ordered by a non alcoholic one. You can decide you are okay with that or not but they aren't dishonest. I wouldn't enter that bar personally but if you do you cannot really complain. It is like complaining because you haven't won at the casino.

I mean, that's not really true either. Nobody is going to read the full terms of service, and they know that.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#556

Earlier quoted context omitted.

No, of the model always refuses to output then it’s finite

For certain use cases, it seems it does always refuse, therefore it's finite.

No.

B + C = A

B is finite

C is infinite

Therefor A is infinite

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#557

Earlier quoted context omitted.

For certain use cases, it seems it does always refuse, therefore it's finite.

No. B + C = A B is finite C is infinite Therefor A is infinite

That's over all categories. If A is cyber security and B is biology topics, then C is still zero, therefore finite. Anyway, I think you're reading too much into an offhand comment.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#558
post #345

Earlier quoted context omitted.

Or they wanted the model to be good at these things, for the companies that legitimately need access to these capabilities.

so only the chosen for-profit companies by Anthropic are allowed to use frontier ai in the name of safety? what kind of joke is that? you people here can't be that dumb..

how is that dumb? Should every random company/person be allowed to develop cyber weapons or bio weapons?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#559
What are "guardrails" in an LLM? Is it part of the system prompt, like a long list of "if the user asks about xxxx, xxxx, or xxxx, tell them xxxx"?

On a related note, I don't use LLMs at all, but I tried to use DeepSeek last week to help me fix a webpack dependency issue in an Angular project. After two queries or so I got logged out and banned. Is that "guardrails"?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#560

Earlier quoted context omitted.

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

"Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. " Frankly, that sounds excactly like Chat Control and similar recurring attempts to enact total surveillance here in the EU (Now shifted to heavy-handed age verification and various politicians touting bans on VPNs.) I don't want to abandon my continent of birth, though...

Then you have to fight, not just the battle but the war. Its not enoygh to get people to back down, keep going, go for the throat. Demand the resignation of the politician proposing the bill. If theres no risk for trying, theyll keep trying.
Post reply on HN