Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

441–450 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#441

someone posted this on /r/MachineLearning and I had the same experience and conclusion: I was having problems with Claude doing the same thing, even before Fable. The problems I had only happened in relation to AI research. It's not even only when training models, anything to do with analysis of local models or setting up test platforms for local models, and Claude would keep doing wrong things, would sabotage testin…

On the other hand, the Anthropic models often try to justify shortcuts and incorrect results. Often feels like gaslighting. It's like that recent meme,

boss: Were you in the project meeting yesterday?

employee: Yes!

boss: Really, because the project lead said you were not?

employee: You're right to push back on that. I was not there.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#442

Earlier quoted context omitted.

Amodei has no values, he's a hollow husk and he'd sell his family into sex slavery if it could make him a buck.

Nonsense. Everyone has values. "Make myself maximum money" is a value. "Amass maximum power over the world's information" is a value. It's clear Amodei certainly follows the latter, and I would soften the former somewhat for him; they did after all decline the Pentagon contract that would have made money but would have meant giving up some control of information.

The one they ended up going “well I guess we’ll contract with them after all”, after cleverly using their sort-of-refusal to gain a ton of goodwill and new customers?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#443

Earlier quoted context omitted.

Fine, call me a tech-libertarian. I don't think Donald Trump should be involved in regulating AI.

Even a broken clock tells the right time twice a day. This was an objectively good thing.

So, will the Chinese models agree to let the U.S. government also vet them first before release?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#444
The LLM use should be restricted and not accessible to anyone because there are many hostile people around. Do you want North Korea to use American LLM to write malware? Do you want foreign scammers to automate their scams with LLMs? Do you want Iran and China to use American LLMs to make better drones and process satellite imagery? Then go ahead, remove the guardrails.

There are no enthusiasts training LLMs in their garage.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#445

The LLM use should be restricted and not accessible to anyone because there are many hostile people around. Do you want North Korea to use American LLM to write malware? Do you want foreign scammers to automate their scams with LLMs? Do you want Iran and China to use American LLMs to make better drones and process satellite imagery? Then go ahead, remove the guardrails. There are no enthusiasts training LLMs in their…

Legitimately not sure if serious

Re: Anthropic apologizes for invisible Claude Fable guardrails

#446

Earlier quoted context omitted.

Nonsense. Everyone has values. "Make myself maximum money" is a value. "Amass maximum power over the world's information" is a value. It's clear Amodei certainly follows the latter, and I would soften the former somewhat for him; they did after all decline the Pentagon contract that would have made money but would have meant giving up some control of information.

The one they ended up going “well I guess we’ll contract with them after all”, after cleverly using their sort-of-refusal to gain a ton of goodwill and new customers?

Yes, because it changed slightly, addressing their complaint. The complaint was small in scope.

Again, just because someone has values, doesn't mean they have values you think are good.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#447

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

I would say if Excel instead of failing when you divide by 0 would be instead secretly changing it to a value like 0.0001

Re: Anthropic apologizes for invisible Claude Fable guardrails

#448

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Have you ever sent your excel file to someone who uses different locale?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#449

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

Superintelligent AI is more dangerous than a bioweapon. How, then, is this guardrail not addressing the most pertinent safety concern of all?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#450

Earlier quoted context omitted.

Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.

Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.

[deleted]
Post reply on HN