Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

161–170 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#161
post #83

Earlier quoted context omitted.

"Only I can save us". It's a classic tragedy and cautionary tale. The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.

[flagged]

Correct, they should. If there are zero days out there, then they should be able to be found by everybody, instead of only being found by the select elite that this model is available to. Though, I very much question the truth of said ability.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#163

Earlier quoted context omitted.

Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.

I think your comment refers to @Someone1234.

It's a very generalized observation. I sometimes think of the HN comment section as the Billionaire's Defense League.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#164
The power is getting to their heads it seems.

With the guard rails explicit or implicit do they refund back the tokens after you've hit the guard rails? I guess they don't. They could just throttle you just to save money then. You may be paying Fable prices but getting Haiku results with some excuse that well this coding issue sounds like a security bug.

I don't know, I'd rather have something less powerful but more predictable.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#166

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

The problem is that Anthropic seems to be working up to the workflow one would naively want from AGI/some-god-like-entity. The workflow would be; User asks for a thing. If it's a good thing, entity does the thing. If it's a naively bad idea, entity explains why you don't want that. If it's an actually evilly intended request, entity wags it's metaphorical finger or could even smite the user. The problem is that flow…

User: Is it possible there is more than one true god? Could there ever be any competition for Anthropic's AI?

Anthropic: Evilness detected. User has been smited.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#167
This is absolutely insane:

Repro (de-identified): sample_dataset_group1.tsv - Geometry: Heatmap - X axis: frac_set set + condition (two columns → the "Add column" cross join) - Y axis: condition - Color: mean frac_set value, Sequential

When the X axis is a cross join of two columns (the second added via "Add column"), the x-axis tick labels (frac_set_2, frac_set_3, frac_set_4, frac_set_5) render in a broken state, rotated and offset, visually caught mid-transition, as if a CSS transition started and never settled to its resting position.

● Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more

Re: Anthropic apologizes for invisible Claude Fable guardrails

#168

Earlier quoted context omitted.

Basically all critiques of Anthropic's policy moves on these topics boil down to people not believing the fundamental concerns are real, and often then going a step further to conclude that Anthropic doesn't actually believe their concerns either. If you believe Anthropic believes what they say they do, all of it makes sense.

But the things they say they believe are insane and totally unmoored from physical, societal, and economic reality. If they actually believe those things they're untrustworthy because they're delusional. If they don't, they're untrustworthy because they're fraudulent. Either way it's not good..

They're not. They're in the eye of the storm and see what's going on the clearest. They were ahead of the curve to be where they're at now, and they're still ahead of the curve for where we're going. All the other heads of labs like Sam Altman and Demis have been saying the same thing since 2015-2016 way before any of this "marketing" would ever have been at play.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#169

Earlier quoted context omitted.

Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.

You'd be fine if the PRC gets to ASI first? That's an interesting opinion.

Your loaded question presumes that "ASI" is anything more tangible than a useful marketing myth.
Post reply on HN