Anthropic apologizes for invisible Claude Fable guardrails
181–190 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#182Earlier quoted context omitted.
Read the actual essay. I cannot possibly imagine how you come to that conclusion unless you're just arguing in bad faith.
No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#183Earlier quoted context omitted.
No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…
This is a pretty reasonable statement and I'm not sure how you could interpret this as "sucking up to the admin."
Re: Anthropic apologizes for invisible Claude Fable guardrails
#184Does "SORRY" fix the deception these models use on the sly?
Does "SORRY" not silently downgrade you to a shittier model without notification?
Does "SORRY" refund your tokens or money?
Im guessing NO to all of those. Standard corporate sorry of "We're sorry youre offended and stupid and gullible".
Re: Anthropic apologizes for invisible Claude Fable guardrails
#185Earlier quoted context omitted.
Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…
[flagged]
Because it’s a threat to ultracapitalist dystopia that they’re tripling down on. The dangers and risk are coming from inside the house.
The danger they care about is the danger to their monopoly, control, and wealth.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#186Re: Anthropic apologizes for invisible Claude Fable guardrails
#187Earlier quoted context omitted.
What is "EA" in this context? I see a lot of people using this initialism.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
Re: Anthropic apologizes for invisible Claude Fable guardrails
#188Re: Anthropic apologizes for invisible Claude Fable guardrails
#189Even on Fable, I'm finding that safeguards can quite easily be surmounted just by incrementally escalating the requests. It's harder than ever to one-shot jailbreaks, but incrementalism still feels like a glaring enough issue to make guardrails just a fig leaf of plausible deniability to the media that they care about "safety."
Re: Anthropic apologizes for invisible Claude Fable guardrails
#190*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails
If by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)