Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

181–190 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#182

Earlier quoted context omitted.

Read the actual essay. I cannot possibly imagine how you come to that conclusion unless you're just arguing in bad faith.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

This is a pretty reasonable statement and I'm not sure how you could interpret this as "sucking up to the admin."

Re: Anthropic apologizes for invisible Claude Fable guardrails

#183

Earlier quoted context omitted.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

This is a pretty reasonable statement and I'm not sure how you could interpret this as "sucking up to the admin."

It's a pretty reasonable statement if you work for Anthropic and are eyeing your stock options nervously and your competitors even more so.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#184
Does "SORRY" fix the invisible garbage guardrails?

Does "SORRY" fix the deception these models use on the sly?

Does "SORRY" not silently downgrade you to a shittier model without notification?

Does "SORRY" refund your tokens or money?

Im guessing NO to all of those. Standard corporate sorry of "We're sorry youre offended and stupid and gullible".

Re: Anthropic apologizes for invisible Claude Fable guardrails

#185
post #95

Earlier quoted context omitted.

Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…

[flagged]

> "Why does a company that cares about the dangers of AI/ASI and x-risk, not want the PRC to catch up to the frontier?"

Because it’s a threat to ultracapitalist dystopia that they’re tripling down on. The dangers and risk are coming from inside the house.

The danger they care about is the danger to their monopoly, control, and wealth.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#187
post #10

Earlier quoted context omitted.

What is "EA" in this context? I see a lot of people using this initialism.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

If you ban women from driving you can eliminate around half the car accidents. Don't you want to reduce car related deaths??

Re: Anthropic apologizes for invisible Claude Fable guardrails

#188
Part of the premise of the article is blatantly wrong. Distillation prevention was always visible. The only invisible safeguard was against frontier model development like development of training pipelines. This doesn't change the general idea that invisible degradation is bad and has been reverted, but the article changes the framing of the original issue from "preventing accelerating AI in the future" to "preventing cheaper AI right now".

Re: Anthropic apologizes for invisible Claude Fable guardrails

#189
> “Visible safeguards can be probed, so they have to be robust, which takes time to get right,” Anthropic wrote.

Even on Fable, I'm finding that safeguards can quite easily be surmounted just by incrementally escalating the requests. It's harder than ever to one-shot jailbreaks, but incrementalism still feels like a glaring enough issue to make guardrails just a fig leaf of plausible deniability to the media that they care about "safety."

Re: Anthropic apologizes for invisible Claude Fable guardrails

#190
post #14

*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails

If by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)

I wasn't buried, it was on the third page after the ToC
Post reply on HN