Earlier quoted context omitted.
"Only I can save us". It's a classic tragedy and cautionary tale. The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.
[flagged]
Anthropic apologizes for invisible Claude Fable guardrails
161–170 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#162Re: Anthropic apologizes for invisible Claude Fable guardrails
#163Earlier quoted context omitted.
Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.
I think your comment refers to @Someone1234.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#164With the guard rails explicit or implicit do they refund back the tokens after you've hit the guard rails? I guess they don't. They could just throttle you just to save money then. You may be paying Fable prices but getting Haiku results with some excuse that well this coding issue sounds like a security bug.
I don't know, I'd rather have something less powerful but more predictable.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#165Re: Anthropic apologizes for invisible Claude Fable guardrails
#166I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
The problem is that Anthropic seems to be working up to the workflow one would naively want from AGI/some-god-like-entity. The workflow would be; User asks for a thing. If it's a good thing, entity does the thing. If it's a naively bad idea, entity explains why you don't want that. If it's an actually evilly intended request, entity wags it's metaphorical finger or could even smite the user. The problem is that flow…
Anthropic: Evilness detected. User has been smited.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#167Repro (de-identified): sample_dataset_group1.tsv - Geometry: Heatmap - X axis: frac_set set + condition (two columns → the "Add column" cross join) - Y axis: condition - Color: mean frac_set value, Sequential
When the X axis is a cross join of two columns (the second added via "Add column"), the x-axis tick labels (frac_set_2, frac_set_3, frac_set_4, frac_set_5) render in a broken state, rotated and offset, visually caught mid-transition, as if a CSS transition started and never settled to its resting position.
● Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more
Re: Anthropic apologizes for invisible Claude Fable guardrails
#168Earlier quoted context omitted.
Basically all critiques of Anthropic's policy moves on these topics boil down to people not believing the fundamental concerns are real, and often then going a step further to conclude that Anthropic doesn't actually believe their concerns either. If you believe Anthropic believes what they say they do, all of it makes sense.
But the things they say they believe are insane and totally unmoored from physical, societal, and economic reality. If they actually believe those things they're untrustworthy because they're delusional. If they don't, they're untrustworthy because they're fraudulent. Either way it's not good..
Re: Anthropic apologizes for invisible Claude Fable guardrails
#169Earlier quoted context omitted.
Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.
You'd be fine if the PRC gets to ASI first? That's an interesting opinion.