Anthropic apologizes for invisible Claude Fable guardrails
111–120 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#112This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…
Stop supporting organizations that don't put humans first. Don't believe a word that anyone says. Lip service is free
Re: Anthropic apologizes for invisible Claude Fable guardrails
#113Earlier quoted context omitted.
Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.
Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#114Earlier quoted context omitted.
Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…
[flagged]
Re: Anthropic apologizes for invisible Claude Fable guardrails
#115Earlier quoted context omitted.
Let's assume that Anthropic believes they're in an arms race to create a potentially dangerous technology, and they believe they're the best ones to win this race. Unlike nuclear weapons, advancing in this arms race requires actually deploying the product over and over again. Deploying the product makes your advancements visible to your competitors. It makes complete sense to try to limit the degree to which that's t…
It's an interesting assumption. The idea behind this with nukes was that we'd like to nuke Germany before they could nuke us. Even after we defeated Germany, we nuked Japan even though they had no possibility of getting their own nukes. The nuclear 'race' was based on the premise that the winner could use it to destroy all other racers (a faulty assumption, see the USSR among others). I will charitably assume Anthrop…
or is it more similar to the Cold War, where there were obviously competitors engaged in the race?
And yes, agreed the equilibrium dynamics for AGI are very different (and far harder to predict) than nukes. That sounds like a good reason to be sure we get there first since presumably any potential advantage wouldn't go to the second or third runner-ups
Re: Anthropic apologizes for invisible Claude Fable guardrails
#116Earlier quoted context omitted.
What are you referring to? The cult belief that they are ushering in a machine god or that they strictly care about making as much money as humanely possibly while ignoring the absolutely destructive impacts these companies have had on society? IMO they are using the cult messaging to distract the public so they take out all the oxygen in the room regarding people that care about the immediate impacts (climate exacer…
"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant. Whether you believe they're creating a machine god is irrelevant. They believe that they are.…
Good to know.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#117Earlier quoted context omitted.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".
But there are also people who just oppose utilitarianism, like G.E.M. Anscombe. For instance, in https://integrityproject.org/wp-content/uploads/2015/07/mr_t..., she seems to grant that dropping the nuclear bombs on Japan was probably good from a utilitarian perspective (because it saved lives overall) and also to grant that bombing campaigns that necessarily entail massive civilian deaths (including, apparently, area bombing German cities) are morally permissible but still to argue that dropping the nuclear bombs was impermissible because it constituted murder ("intentionally" killing the innocent). But this kind of distinction, which I think is what actual anti-utilitarianism must come to, is hard to even consistently maintain, and I suppose many HN readers would find the effort quixotic.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#118Re: Anthropic apologizes for invisible Claude Fable guardrails
#119They relied on trust that they were providing the service they were being paid for. That trust was blown, and an "oops, lets undo that" does not regain trust. It would be prudent to assume the invisible guardraild are possibly in play for all future Clause use, Fable or otherwise.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#120Earlier quoted context omitted.
The hidden safeguard was not against distilling, it was against "frontier" ML research with no indication whatsoever of what "frontier" might mean, but possibly even including research into model safety or alignment. That amounts to deliberately boobytrapping research across an entire legit academic field, which is ridiculously unaligned behavior.
This is the same as saying "well some unaligned countries will use refined nuclear material for energy, too!" lmao. The vast majority of frontier research is about how to build better models, not about alignment.