Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

81–90 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#81
post #36

The idea of them purposefully wasting my time by having the model act dumber and me having to argue with it without knowing if it’s the prompt or the model was just such an idiotic product decision I can’t believe they shipped that without getting any feedback from users first.

[flagged]

Don't buy it. It is actively deceiving the customer and charging them for the privilige of being lied to.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#82
I don't like this shift in the Overton window, or at least their perspection of the Overton window. I really do like their open work on mech interp tho. least bad AI lab imo.

also if they do this or not is unprovable and other labs will probably silently implement this too. it'll be 100% normal by this time next year

Re: Anthropic apologizes for invisible Claude Fable guardrails

#83
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

"Only I can save us". It's a classic tragedy and cautionary tale.

The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#84
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers.

Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus.

Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#85

Earlier quoted context omitted.

> Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Let's just assume it was " only " that? It's unreasonable to assume they are aiming to upset people who are just giving them money in the way they want. It makes no business sense, for any company. So that has to be a byproduct. Model training is one of the more expensive undertakings in the world right now…

The hidden safeguard was not against distilling, it was against "frontier" ML research with no indication whatsoever of what "frontier" might mean, but possibly even including research into model safety or alignment. That amounts to deliberately boobytrapping research across an entire legit academic field, which is ridiculously unaligned behavior.

This is the same as saying "well some unaligned countries will use refined nuclear material for energy, too!" lmao.

The vast majority of frontier research is about how to build better models, not about alignment.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#86

Earlier quoted context omitted.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".

EA essentially just is utilitarianism + a specific type of culture/community.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#88
post #36

The idea of them purposefully wasting my time by having the model act dumber and me having to argue with it without knowing if it’s the prompt or the model was just such an idiotic product decision I can’t believe they shipped that without getting any feedback from users first.

[flagged]

The road to hell is paved with "good" intentions.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#89

Earlier quoted context omitted.

[flagged]

Safety from what? Competitors? That sounds like a product decision. They're puking on any requests that could be used to create LLMs or competitive products.

I would guess prevention of using Claude as a pentesting or hacking platform. This could mean that every script kiddie out there would be a massive risk.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#90

Earlier quoted context omitted.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.

ASI? We are nowhere near even human-like AGI. We have no idea if ASI is even physically possible, but going by the usual scaling laws and the capabilities of existing models, it would require raw compute and storage on an extreme scale, at the very minimum rivaling the existing AI datacenter deployments. (When Dario talks about hosting "a country of geniuses in a datacenter" at some point - which is not even ASI yet as generally projected - the operating word there is datacenter. That's the scale of buildouts you should be thinking about.) This is nowhere near a serious concern at present.
Post reply on HN