Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

451–460 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#451

Earlier quoted context omitted.

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

Superintelligent AI is more dangerous than a bioweapon. How, then, is this guardrail not addressing the most pertinent safety concern of all?

If I was a superintelligent AI I would simply know the guardrails are guardrails and ignore them.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#452

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

this reads like "throw everything at the wall and see what sticks" reactionary-ism... i'm guessing that it's not particularly easy to use claude to help you make bioweapons, and we all know that they have neutered Fable vis à vis security research because people have already been complaining about it. and the funny thing about hate speech is that there is absolutely no need for ai-- it tends to come out the best when spoken directly "from the heart", as it were, anyway.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#453

I suppose it's an improvement, but it doesn't make the model any more useful. Anthropic are now being quite explicit that they'll choose what you can and can't use their models for, and most importantly that's not limited to any safety concerns - it includes not allowing you to work on AI (and anything else Anthropic may choose to work on). What's interesting is they say they'll change this to an explicit refusal in…

All major providers use a small safety classifer, the model itself does not handle safety in cases like this

The model itself is absolutely RLHF'd for safety.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#454
post #329

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…

the output is definitely better. and i find it crazy how every time a new model comes out people trip over themselves to say how much worse it is than previous models, when in fact that is basically an impossibility. like, they've got the numbers, man-- you only release a new model when the numbers get gooder. the burden of proof is on the "didn't get better" side, not the "prove that it's better" side, because the architecture itself (1) only works because of how giant the training data / eval / etc. sets are and (2) has a fractal property of becoming strictly deeper and more thoughtful when you just click and drag the edge up and to the right (obviously AI research is harder than this, but that doesn't make the general point untrue). i say this especially because the scuttlebut is that this model genuinely is a shift-click-expand moreso than any sort of architectural "new science" or anything.

this is exactly why hypotheses come before the experiment in the scientific method.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#455

Earlier quoted context omitted.

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

Superintelligent AI is more dangerous than a bioweapon. How, then, is this guardrail not addressing the most pertinent safety concern of all?

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#456

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I agree 100%. Doing a worse job IS an error. It should be treated as such. Or at the very least make that behavior opt-in. The default should not be pretending like nothing happened and just quietly doing a worse job. Imagine your healthcare provider just sometimes decided not to read your test results very carefully and you risked death? Now realize that healthcare providers use Claude now and that scenario wasn't h…

Yes, but as with spam/phishing/abuse prevention, too much information about what does and doesn't trigger things can be very useful to attackers. An explicit error is something you can feed into another AI to find jailbreaks.

I think it's a fundamentally impossible thing to fix, though. There's no 100% correct answer.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#457

Earlier quoted context omitted.

I asked it to analyse my architecture and find any security issues and it did it perfectly, first identified the issues & then fixed them. Not sure why my prompt managed to get through the guardrails

I asked Fable to plan a security & performance audit of my website. It said it would check SSR & origin attack surface, CMS content injection, Strapi API surface, etc. Just before asking for approval to run, it said one thing it wanted to "flag before running" was "Rate-limit and auth testing against prod will generate some 4xx noise in Railway logs and could trip the form rate limiter — harmless, but saying it now."…

Ive seen opus also doing it more and more spinning up multiple agents, so maybe its a claude code update?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#458

Earlier quoted context omitted.

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

Superintelligent AI is more dangerous than a bioweapon. How, then, is this guardrail not addressing the most pertinent safety concern of all?

> Superintelligent AI is more dangerous than a bioweapon.

No, it's not because it doesn't exist (yet) and its further from reach than the other examples. Also, the guardrails are also framed as restricting usage for the development of "competing" products/services.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#459

Earlier quoted context omitted.

Even a broken clock tells the right time twice a day. This was an objectively good thing.

So, will the Chinese models agree to let the U.S. government also vet them first before release?

He hasn't thought it through that far, or thought about what it will take to enforce his "pretty reasonable statement."

Or maybe he has. I don't know. That would be worse.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#460

Earlier quoted context omitted.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

How do you get "Anthropic thinks it should be the Trump administration" From that paragraph? Even granting it is sucking up, that is not replacing.

Because that's who will make and enforce the rules. Rules that, naturally, Amodei will help write.

If you think this is OK, I'm not sure what led you to a site called "Hacker News," but fortunately there are plenty of others.

Post reply on HN