Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

461–470 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#461

Earlier quoted context omitted.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

I agree with your sentiment but not your conclusion. They don't want this administration specifically to have gatekeeping authority, what they want is any administration to say that they are gatekeeping, so that they can regulate the competition out of existence. Of course the actual checks and balances will be near pointless in effect, but expensive to implement nonetheless.

Of course the actual checks and balances will be near pointless in effect

My concern is that they won't be pointless in effect. Make no mistake: if Amodei has his way, possession of unvetted model weights will be treated like possession of CSAM is today. And at the same time Amodei calls for that, others are calling for the deployment of technical measures that will make it easier to enforce such laws.

All to the sound of thunderous applause on "Hacker News."

Re: Anthropic apologizes for invisible Claude Fable guardrails

#462

Earlier quoted context omitted.

It's not just America. The main secret is out of the bag. If it wasn't Anthropic, it would be another company/nation state. Sure they could obtain, and with not.money or leverage, complain about data centers at local rallies, or they can be in the game, and hopefully steer it. It's going to happen with or without any one company or country. The secret it out, and it's unstoppable without complete societal breakdown..…

> It's not just America. I'll mention again the nuclear analogy. It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies. > The main secret is out of the bag. This is not something you can do in a shed with a handful of GPUs just because you know "the main secret". To build something like Mythos you need tens o…

Yeah, I think the nuclear analogy fails, honestly. The bomb does one thing: Destroy. AI can build, selectively infect, selectively manipulate. It is a vastly useful tool.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#463

Earlier quoted context omitted.

> I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviously. I think you're reading some subtext into my comment that I didn't intend. Knowing myself, I assume the scare quotes there are just a bit of casual irony re: the insanely high stakes here. The word "concerns" as use…

> You can, in fact, opt out. You can opt out and do your damndest to stop what's happening, throw every cent you have at it, bend any ear that will listen, make use of the fact that your voice (as Anthropic leadership) has some meaningful weight. There are billions of people who have opted out of playing the game. Has the game stopped? Has any game stopped because the people not playing it decided that it ought to? O…

(I wrote a longer comment originally, but I think it would have fallen on deaf ears.)

> The only variable of disagreement is around AI doom.

The source of our disagreement seems to be your belief that somebody can either a) believe "AI doom" is inevitable, or b) not believe it's possible. This is an obvious false dichotomy that's stunting your ability to engage effectively with what I've written, and also stunting your ability to understand the broader landscape of the issue. You are projecting this dichotomy onto everybody involved and understanding their behavior in that way, which is leading you to make other reductive and honestly bizarre claims—like, for instance, the idea that a sudden change of course from Dario Amodei at this moment in time would be broadly perceived as somebody who was already losing the race cashing out his chips. If you really do believe that, I have to assume it's because you're modeling your hypothetical observers as falling into one of your two extreme mindsets and assuming Dario, being a smart guy who knows a lot, thinks the same way. I believe it is—yes, I'm ready for the dopamine hit—naive to assume all people or even most people fall into one of these two camps. Naive, at best. Yours is a self-limiting framework for thinking about this stuff.

I encourage you to broaden your thinking and engage in less projection and ad hominem stuff in discussions like this. I probably won't reply to whatever you post next unless you can do a better job writing a substantive reply to what I've written here.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#464

Earlier quoted context omitted.

How do you get "Anthropic thinks it should be the Trump administration" From that paragraph? Even granting it is sucking up, that is not replacing.

Because that's who will make and enforce the rules. Rules that, naturally, Amodei will help write. If you think this is OK, I'm not sure what led you to a site called "Hacker News," but fortunately there are plenty of others.

No, it is not OK. But also, not what you quoted says.

Not sure who you are arguing with really. There seems to be a few logical leaps in between each response. I also didn't say anything like that.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#465

Earlier quoted context omitted.

Because that's who will make and enforce the rules. Rules that, naturally, Amodei will help write. If you think this is OK, I'm not sure what led you to a site called "Hacker News," but fortunately there are plenty of others.

No, it is not OK. But also, not what you quoted says. Not sure who you are arguing with really. There seems to be a few logical leaps in between each response. I also didn't say anything like that.

You asked, I answered. The direct and immediate effect of Amodei getting what he asks for in that essay will be to empower the Trump administration to approve model releases.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#466

Earlier quoted context omitted.

> It's not just America. I'll mention again the nuclear analogy. It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies. > The main secret is out of the bag. This is not something you can do in a shed with a handful of GPUs just because you know "the main secret". To build something like Mythos you need tens o…

Yeah, I think the nuclear analogy fails, honestly. The bomb does one thing: Destroy. AI can build, selectively infect, selectively manipulate. It is a vastly useful tool.

This is a very imprecise way to think about it.

What is the difference? It's easier to make money with the AI you get at each incremental step toward potentially destroying human civilization (though, of course, it's debatable whether these companies really are making money as such).

So what? You are implicitly arguing that human civilization will be unable to resist engaging in a large-scale, coordinated effort to destroy itself, just to make a few bucks along the way. Is this true? I don't know. The point of the nuclear analogy is that we have previously shown that we can, under certain conditions, put the eschaton back on the shelf for some period of time, despite very real pressure to take more incremental steps toward doom. "But AI can write code" is not a refutation of the possibility that we could take a more measured approach to AI development.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#467

Earlier quoted context omitted.

Yeah, I think the nuclear analogy fails, honestly. The bomb does one thing: Destroy. AI can build, selectively infect, selectively manipulate. It is a vastly useful tool.

This is a very imprecise way to think about it. What is the difference? It's easier to make money with the AI you get at each incremental step toward potentially destroying human civilization (though, of course, it's debatable whether these companies really are making money as such). So what? You are implicitly arguing that human civilization will be unable to resist engaging in a large-scale, coordinated effort to d…

I never claimed to be precise, ha ha. But let's not lose sight that each of the great powers agreed to limits on the number of nukes only after making enough to collectively destroy each other several times over. They stopped exactly when building more did not gain them any more security or advantage over their adversaries. Not a moment before that.

There may or may not be such a point with AI: A point at which ever smarter machines provides no marginal benefit to security. If that happens, I do expect agreements, should any one of us still exist.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#468

Earlier quoted context omitted.

Amodei has no values, he's a hollow husk and he'd sell his family into sex slavery if it could make him a buck.

Nonsense. Everyone has values. "Make myself maximum money" is a value. "Amass maximum power over the world's information" is a value. It's clear Amodei certainly follows the latter, and I would soften the former somewhat for him; they did after all decline the Pentagon contract that would have made money but would have meant giving up some control of information.

Those aren't values. Maybe goals or motivations but not values in any conceivable way, shape or form. This site is full of pod people I swear.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#469

Earlier quoted context omitted.

Everyone that isn't a bitter cynic must be a shill.

I’ve noticed that too many HN folks seem to think that cynicism makes them more intelligent. I think it must be some kind of insecurity, about not wanting to be seen as naive or something. It’s pretty sad though, I wonder how some of these people find any peace or joy in their lives.

Believing they have any interests other than theirs at heart is like believing the stripper is really in love with you. That's not cynicism, that's just common sense.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#470
post #415

Earlier quoted context omitted.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

I may be naive, but I have the feeling that "I will arbitrarily set numbers on things and call it impartial" is... weird at best. I understand how one may wonder if there was a way to do that, but it feels insane to me that one would actually conclude that "yes, it is possible". We have examples everywhere showing that it is generally impossible to define a metric that correctly represents the underlying concept we w…

To quote notorious effective altruist Scott Alexander:

> Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”.

https://slatestarcodex.com/2017/04/07/yes-we-have-noticed-th...

Post reply on HN