Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

111–120 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#112
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Corporation cannot help but act this way. They are too big. The pressures for profit are all that matters. That is the priority. It doesn't matter what colorful words they put on the paper to make you feel better. Look at the "green" movement 20 years ago. All talk and no action.

Stop supporting organizations that don't put humans first. Don't believe a word that anyone says. Lip service is free

Re: Anthropic apologizes for invisible Claude Fable guardrails

#113

Earlier quoted context omitted.

Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.

Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#114
post #95

Earlier quoted context omitted.

Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…

[flagged]

I don’t think they’re mutually exclusive. It’s a business selling a product that isn’t yet profitable, not a public advocacy organization.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#115

Earlier quoted context omitted.

Let's assume that Anthropic believes they're in an arms race to create a potentially dangerous technology, and they believe they're the best ones to win this race. Unlike nuclear weapons, advancing in this arms race requires actually deploying the product over and over again. Deploying the product makes your advancements visible to your competitors. It makes complete sense to try to limit the degree to which that's t…

It's an interesting assumption. The idea behind this with nukes was that we'd like to nuke Germany before they could nuke us. Even after we defeated Germany, we nuked Japan even though they had no possibility of getting their own nukes. The nuclear 'race' was based on the premise that the winner could use it to destroy all other racers (a faulty assumption, see the USSR among others). I will charitably assume Anthrop…

Do you believe the current situation is more akin to the race to the first nukes, where no one could know for sure the other competitors were even racing...

or is it more similar to the Cold War, where there were obviously competitors engaged in the race?

And yes, agreed the equilibrium dynamics for AGI are very different (and far harder to predict) than nukes. That sounds like a good reason to be sure we get there first since presumably any potential advantage wouldn't go to the second or third runner-ups

Re: Anthropic apologizes for invisible Claude Fable guardrails

#116
post #55

Earlier quoted context omitted.

What are you referring to? The cult belief that they are ushering in a machine god or that they strictly care about making as much money as humanely possibly while ignoring the absolutely destructive impacts these companies have had on society? IMO they are using the cult messaging to distract the public so they take out all the oxygen in the room regarding people that care about the immediate impacts (climate exacer…

"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant. Whether you believe they're creating a machine god is irrelevant. They believe that they are.…

Oh okay, they're all just legit crazy and are allowed to poison the environment, murder teenagers, and ruin the material lives of millions for fantasy level delusions.

Good to know.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#117

Earlier quoted context omitted.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".

It's important to separate objections to utilitarianism from the obvious fact that it can very be hard to correctly apply the utilitarian calculus. It's partly because of this difficulty that most classical utilitarians thought that people should generally follow commonsense morality and not try to directly apply the utilitarian calculus (which then led to the charge of paternalism and teaching one morality to the masses and another to a supposed elite).

But there are also people who just oppose utilitarianism, like G.E.M. Anscombe. For instance, in https://integrityproject.org/wp-content/uploads/2015/07/mr_t..., she seems to grant that dropping the nuclear bombs on Japan was probably good from a utilitarian perspective (because it saved lives overall) and also to grant that bombing campaigns that necessarily entail massive civilian deaths (including, apparently, area bombing German cities) are morally permissible but still to argue that dropping the nuclear bombs was impermissible because it constituted murder ("intentionally" killing the innocent). But this kind of distinction, which I think is what actual anti-utilitarianism must come to, is hard to even consistently maintain, and I suppose many HN readers would find the effort quixotic.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#119
I don't think they can convince me they have actually reversed course on this. Its invisible so we wouldn't know if they kept on doing it secretly. It required building out technical capability which is unlikely to remain forever unused while conveniently available to them.

They relied on trust that they were providing the service they were being paid for. That trust was blown, and an "oops, lets undo that" does not regain trust. It would be prudent to assume the invisible guardraild are possibly in play for all future Clause use, Fable or otherwise.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#120

Earlier quoted context omitted.

The hidden safeguard was not against distilling, it was against "frontier" ML research with no indication whatsoever of what "frontier" might mean, but possibly even including research into model safety or alignment. That amounts to deliberately boobytrapping research across an entire legit academic field, which is ridiculously unaligned behavior.

This is the same as saying "well some unaligned countries will use refined nuclear material for energy, too!" lmao. The vast majority of frontier research is about how to build better models, not about alignment.

And as a matter of fact, there's a lot of meaningful research into how to have different sorts of nuclear material that might be usable for power production but not hidden malicious development. That's the closest analog to "safety" and "alignment" in your scenario.
Post reply on HN