Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

191–200 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#191

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word.

Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up to their product. They care less about hindering bioweapons, malware, and hate speech than they do free market competition.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#192

Earlier quoted context omitted.

Do you believe the current situation is more akin to the race to the first nukes , where no one could know for sure the other competitors were even racing... or is it more similar to the Cold War, where there were obviously competitors engaged in the race? And yes, agreed the equilibrium dynamics for AGI are very different (and far harder to predict) than nukes. That sounds like a good reason to be sure we get there…

I can't really say I see a similarity to either the Manhattan Project or the Cold War. I don't see how one could apply either massive retaliation or MAD. These are private companies, they are not vested with the necessary authority to destroy anything. Even if they had it, they couldn't. You can't destroy China, they have 1.4B people, nukes, and a large part of the world's manufacturing. So multiple organizations wan…

You think "arms race" is a dynamic that only applies to literal arms?

"Ability to literally destroy the other entity" is not a necessary or even typical feature of arms races.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#193

Earlier quoted context omitted.

They're not. They're in the eye of the storm and see what's going on the clearest. They were ahead of the curve to be where they're at now, and they're still ahead of the curve for where we're going. All the other heads of labs like Sam Altman and Demis have been saying the same thing since 2015-2016 way before any of this "marketing" would ever have been at play.

There's a simpler explanation that fits the data better: they're lying. Generally, in the past when tech companies have made outlandish claims that were not backed by evidence, they're later found out to have lied. This is an ancient pattern going back to the dotcom era and before, but for recent examples you need only look back a few years to the web3 era. If they're not lying, they can show it by producing the resu…

What data does "they're lying" fit better than "they're earnest?"

> If they're not lying, they can show it by producing the results they claim. Until then, they're probably just lying

Brilliant framework: Anyone making claims about the future is not just speculating, not just wrong, but they are lying.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#194

Earlier quoted context omitted.

I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".

It's important to separate objections to utilitarianism from the obvious fact that it can very be hard to correctly apply the utilitarian calculus. It's partly because of this difficulty that most classical utilitarians thought that people should generally follow commonsense morality and not try to directly apply the utilitarian calculus (which then led to the charge of paternalism and teaching one morality to the ma…

The first half of your answer presupposes some platonic utilitarian calculus that, if it were applied correctly, would yield moral outcomes. This is very hard to believe. If I look at notable/well-known examples of EA-affiliated people, it is hard to skip by members such as SBF. Did he correctly apply the utilitarian calculus?

It is relatively easy to take the proceeds of a massive fraud, buy a relatively small (as a percentage of the fraud) $ amount of mosquito nets, and save more lives than the lives impacted by your massive theft. Is this a correct application of the utilitarian calculus? What sort of data would we need a priori to do this calculation "correctly"? Do you think he had a careful estimate of the suicide rate of victims of ponzi schemes before perpetuating the fraud, or would any suicide rate have made the decision net [pun intended] moral, as any such victim of fraud would lead to >> 1 net purchased (so you would almost always net save lives).

The above is of course snarky. It is also a best-effort way of analyzing a notable utilitarian's actions. I do not think it would be difficult at all to use this type of argument to argue that SBF's actions net raised utility in the world. If only we all would become fraudsters, then we could truly live in Omelas --- a notable utilitarian paradise.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#195
ITT a surprising lack of perspective on the fact that despite the breathless pace of the singularity, people are still necessarily figuring things out as we go and we are well off the map.

Here there be monsters, and we don't have any real way of evaluating risk; and the leverage provided by tools already available affords systemic and even existential risk in a way no one—least of all an industry committed to shareholder value—has had to navigate, let alone with a million backseat drivers each with their own substack and brand to build.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#196

Earlier quoted context omitted.

They performed famously well at FTX.

Guess FTX disproved the concept of giving to effective charities, time to start donating to my church again.

What FTX decisively disproved was the idea that people's origin stories involving apparently sincere desire to do good in the world and them constantly broadcasting that should be used as a reason to unquestioningly trust them when their notion of greater good happens to align perfectly with them accumulating enormous quantities of wealth and power. (and Sam, bless him, originally wanted to help animals rather than own the machine god. And probably sincerely believed he was going to do great things for humanity from all the misappropriated funds he was definitely going to win back against a backdrop of EAs and VCs queueing up to glaze him and his commitment to the greater good)

I don't think people are objecting to the EA idea that some charities are more evidence based than others so much as the distinctly EA idea that it would be more effective still to donate to charities like OpenAI

Re: Anthropic apologizes for invisible Claude Fable guardrails

#197

Earlier quoted context omitted.

I can't really say I see a similarity to either the Manhattan Project or the Cold War. I don't see how one could apply either massive retaliation or MAD. These are private companies, they are not vested with the necessary authority to destroy anything. Even if they had it, they couldn't. You can't destroy China, they have 1.4B people, nukes, and a large part of the world's manufacturing. So multiple organizations wan…

You think "arms race" is a dynamic that only applies to literal arms? "Ability to literally destroy the other entity" is not a necessary or even typical feature of arms races.

Well it's difficult to argue against something that was never specifically stated. If someone is able to state specifically how this is an arms race in any other way than that it's a race at all then I'm happy to have that conversation.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#198
post #107

Earlier quoted context omitted.

How does US regulatory capture do anything to impede PRC's advance?

Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#199
post #136
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Yes, that is basically the plan. It's based on the belief that unfettered AI would let anyone be a supervillain and destroy the world. There are enough would-be supervillains out there, but they rarely get far because they can't get teams of smart people to build doomsday machines for them. So the AI has to not let anyone do evil with it. Unfortunately, that won't feel very much like freedom.

It sounds like you might not agree with that belief.

While I don't agree with their actions here, I do think there's sufficient reason to hold that belief.

On some fronts (e.g. security, on which you've experienced more than me), I think there are surmountable challenges. But on other fronts (e.g. bio), a single errant actor could reasonably kill millions or billions of people with sufficiently powerful AI. We don't have good defenses here, and those actors do exist.

I still don't agree with these actions, but I do think I agree with their assumptions.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#200

Earlier quoted context omitted.

You think "arms race" is a dynamic that only applies to literal arms? "Ability to literally destroy the other entity" is not a necessary or even typical feature of arms races.

Well it's difficult to argue against something that was never specifically stated. If someone is able to state specifically how this is an arms race in any other way than that it's a race at all then I'm happy to have that conversation.

"Arms race" is the term used colloquially to describe the dynamic that emerges in "winner-take-all" markets.

It seems that the frontier labs believe they're participants in a winner-take-all market. Therefore they're in "an arms race."

Winner-take-all markets do not require that the winner literally destroys the losers, but only that the winner enjoys disproportionate returns compared to their actual superiority.

Whether or not this is actually true is TBD, but I think you're naive to think the frontier labs do not believe this to be true.

Post reply on HN