Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

431–440 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#431

Earlier quoted context omitted.

It's not just America. The main secret is out of the bag. If it wasn't Anthropic, it would be another company/nation state. Sure they could obtain, and with not.money or leverage, complain about data centers at local rallies, or they can be in the game, and hopefully steer it. It's going to happen with or without any one company or country. The secret it out, and it's unstoppable without complete societal breakdown..…

> It's not just America. I'll mention again the nuclear analogy. It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies. > The main secret is out of the bag. This is not something you can do in a shed with a handful of GPUs just because you know "the main secret". To build something like Mythos you need tens o…

> It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies.

You were literally just criticizing Anthropic as disingenuous for begging for this. Or is your position that people other than those near the front of the race can agree to limit development? And if so: provide evidence.

Note also a key ingredient that makes nuclear non-proliferation possible is that they're pretty much useless weapons. There is no smaller order nuke that's dramatically more useful than a large conventional weapon. That's not true of AI models, which appear to be monotonically useful as they become more powerful.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#432

Earlier quoted context omitted.

How do you think the Qwen and MiniMax models perform so similarly to Anthropic frontier models? What is your take then?

Well Anthropic did not ask for permission before they distilled copyrighted material. At least the Chinese have the decency of giving back the model weights and not put BS censorship because “it’s too dangerous”.

Ask DeepSeek about Tianamen Square and see what happens. The Chinese models have censorship too.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#433

Earlier quoted context omitted.

> (which depend, by the way, on accelerating the world toward those hypothetical "concerning" scenarios as fast as possible) Yes, this dynamic is exactly the one that anyone who's concerned about AI is concerned about. I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviousl…

> I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviously. I think you're reading some subtext into my comment that I didn't intend. Knowing myself, I assume the scare quotes there are just a bit of casual irony re: the insanely high stakes here. The word "concerns" as use…

> You can, in fact, opt out. You can opt out and do your damndest to stop what's happening, throw every cent you have at it, bend any ear that will listen, make use of the fact that your voice (as Anthropic leadership) has some meaningful weight.

There are billions of people who have opted out of playing the game. Has the game stopped? Has any game stopped because the people not playing it decided that it ought to? Only with government intervention, which is exactly what you just criticized Anthropic for being disingenuous for requesting.

Is your position that they should just be smart bloggers asking for regulation, instead of the preeminent lab asking for regulation, and that would be either more ethical or more effective? If it's less effective, isn't it de facto less ethical?

What say you about the thousands of smart bloggers asking for regulation who are ignored every single day and have no tools besides their blogs to steer the outcome?

> burn your position in the race to show just how fucking serious you are.

This is incredibly naive. Literally no one who is unconvinced of AI doom would be convinced by this... because they already don't believe the premise. Such a gesture would be readily explained away as "you were losing the race," or "you got rich enough already." This is the attitude when any individual opts out of participation (see: Hinton) and it's ridiculous to assume it'd be different if an entire company did it.

Not to mention, that an entire company can't do it. These companies have boards of directors. They are accountable to shareholders. A CEO who wanted to do this would simply be fired and the company would carry on. This is one of the key components of the trap. Large companies are not under the control of people but of incentives. They are literally deliberately designed not to be under the control of individuals –– to be immune to exactly the type of behavior you think is possible.

And yes, nuclear weapons are analogous to AI in the arms race dynamic to create and proliferate them. They are probably not analogous in there exists a stable equilibrium in nuclear weapons due to "accidents" of their nature. There need not be a similar equilibrium among competitive AIs.

----

And yes, your comment lands in exactly the category I mentioned. You do not believe the AI doom fears, so the behavior looks one way. I do believe in the AI doom concerns, so the behavior looks another way. This applies equally to yesterday's actions, today's actions, and tomorrow's, including some hypothetical honorable self-immolation to slow progress: I would see that as concordant with AI doom concerns, you would not. You would find it "hard not to be cynical" about the fact they already earned a ton of money, maybe was losing a bit of ground in the race, so on and so forth. This is plainly obvious to anyone who has had to converse with no-doomers, who can only analyze other people's behaviors under their own belief system, so it won't happen.

The only variable of disagreement is around AI doom.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#434

Earlier quoted context omitted.

"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant. Whether you believe they're creating a machine god is irrelevant. They believe that they are.…

Sometimes governments have to deal with the weapons made by their enemies and that gets them stuck in an arms race. Companies don't have to do that. If they're getting into actually dangerous territory, they can stop as soon as they want to.

If you believe in the AI doom scenario then yes, you do need to do that. Because it's very important that your "less ethical" and "less good" competitors do not get to the machine god first.

If you don't believe that, or you don't believe that the frontier labs believe that, then sure, it makes no sense. But they probably do. The people at these companies literally dedicated their lives to building this specific thing that, up until people had to make tradeoffs between "that looks risky" and "that looks useful", virtually everyone agreed would be a dangerous technology.

What apparently many people on HN failed to appreciate is that the thing that makes it dangerous is the fact that it grows in utility.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#436

Earlier quoted context omitted.

This is like the most milquetoast stance in the AI safety community. It's great the Trump admin did something, no one expected them to, and they should have done more. Very powerful tools released to the public should be regulated for safety. That is "pretty reasonable" to most people (except the tech-libertarian crowd).

Fine, call me a tech-libertarian. I don't think Donald Trump should be involved in regulating AI.

Even a broken clock tells the right time twice a day. This was an objectively good thing.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#437

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

The "look", of course, is completely bullshit. Release the model, give licensing terms, sue the ever living daylights of anyone who's hosting it without agreeing to those daylights, and move on. This vertical integration shit that we're all enamored with is bullshit. Even Amazon has their own vans inside of UPS being their own thing? No wonder stepmom porn is on the rise.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#438

Earlier quoted context omitted.

Read the actual essay. I cannot possibly imagine how you come to that conclusion unless you're just arguing in bad faith.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

How do you get "Anthropic thinks it should be the Trump administration"

From that paragraph?

Even granting it is sucking up, that is not replacing.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#439

Earlier quoted context omitted.

"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant. Whether you believe they're creating a machine god is irrelevant. They believe that they are.…

Sometimes governments have to deal with the weapons made by their enemies and that gets them stuck in an arms race. Companies don't have to do that. If they're getting into actually dangerous territory, they can stop as soon as they want to.

Joking? Companies absolutely do get into arms races.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#440

Earlier quoted context omitted.

"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant. Whether you believe they're creating a machine god is irrelevant. They believe that they are.…

A lot of people would prefer nuclear deproliferation over building more nukes. Arms races always work out great for arms dealers. Less so for the average Joe.

Yes, totally. But those people's preferences don't matter. What matters is the core competitive dynamic.
Post reply on HN