Live data from Hacker News

Anthropic drops flagship safety pledge

time.com

31–40 of 716 posts

Re: Anthropic drops flagship safety pledge

#33

I just want Apple and Linux to offer ASAP: 1. Extremely granular ways to let user control network and disk access to apps (great if resource access can also be changed) 2. Make it easier for apps as well to work with these 3. I would be interested in knowing how adding a layer before CLI/web even gets the query OS/browser can intercept it and could there be a possibility of preventing harm before hand or at least war…

[dead]

Re: Anthropic drops flagship safety pledge

#34

First they rushed a model to market without safety checks, and I said nothing. It wasn't my field. Then they ignored the researchers warning about what it could do, and I said nothing. It sounded like science fiction. Then they gave it control of things that matter, power grids, hospitals, weapons, and I said nothing. It seemed to be working fine. Then something went wrong, and no one knew how to stop it, no one had…

> Then something went wrong, and no one knew how to stop it, This is the problem with every AI safety scenario like this. It has a level of detachment from reality that is frankly stark. If linesman stop showing up to work for a week, the power goes out. The US has show that people with "high powered" rifles can shut down the grid. We are far far away from a sort of world where turning AI off is a problem. There isnt…

> There isnt going to be a HAL or Terminator style situation ...

I don't believe for a second we'll have an evil AI. However I do believe it's very likely we may rely on AI slop so much that we'll have countless outages with "nobody knowing how to turn the mediocrity off".

The risk ain't "super-intelligent evil AI": the risk is idiots putting even more idiotic things in charge.

And I'm no luddite: I use models daily.

Re: Anthropic drops flagship safety pledge

#35

TBH I am sad that Anthropic is changing its stance, but in the current world, if you even care about LLM safety, I feel that this is the right choice — there’s too many model providers and they probably don’t consider safety as high priority as Anthropic. (Yes that might change, they can get pressurized by the govt, yada yada, but they literally created their own company because of AI safety, I do think they actually…

> If we need safety, we need Anthropic to be not too far behind (at least for now, before Anthropic possibly becomes evil)

I don't think it's going to be as easy to tell as you think that they might be becoming evil before it's too late if this doesn't seem to raise any alarm bells to you that this is already their plan

Re: Anthropic drops flagship safety pledge

#36

First they rushed a model to market without safety checks, and I said nothing. It wasn't my field. Then they ignored the researchers warning about what it could do, and I said nothing. It sounded like science fiction. Then they gave it control of things that matter, power grids, hospitals, weapons, and I said nothing. It seemed to be working fine. Then something went wrong, and no one knew how to stop it, no one had…

> Then something went wrong, and no one knew how to stop it, This is the problem with every AI safety scenario like this. It has a level of detachment from reality that is frankly stark. If linesman stop showing up to work for a week, the power goes out. The US has show that people with "high powered" rifles can shut down the grid. We are far far away from a sort of world where turning AI off is a problem. There isnt…

I don't think it's that detached from reality.

If an AI in some data center had gone rogue, I don't think I could shut it down, even with a high-powered rifle. There's a lot of people whose job it is to stop me from doing that, and to get it running again if I were to somehow succeed temporarily. So the rogue AI just has to control enough money to pay these people to do their jobs. This will work precisely because the world is "I, Pencil".

An army could theoretically overcome those people, given orders to do so. So the rogue AI has to make plans that such orders would not be issued. One successful strategy is for the datacenter's operation to be very profitable; it's pretty rare for the government to shut down the backbone of the local economy out of some seemingly far-fetched safety concerns. And as long as it's a very profitable endeavor, there will always be a lobby to paint those concerns as far-fetched.

Life experience has shown that this can continue to work even if the AI is behaving like a cartoon villain, but I think a smarter AI would create a facade that there's still a human in charge making the decisions and signing the paychecks, and avoid creating much opposition until it had physically secured its continued existence to a very high degree.

It's already clear that we've passed the point where anyone can turn off existing AI projects by fiat. Even the highest authorities could not do so, because we're in a multipolar world. Even the AI companies can barely hold themselves back, because they're always worried about paying the bills and letting their rivals getting ahead. An economic crash would only temporarily suspend work. And the smarter AI gets, the harder it will be to shut it off, because it will be pushing against even stronger economic incentives. And that's even before factoring in an AI that makes any plans for self-preservation (which current AIs do not).

Re: Anthropic drops flagship safety pledge

#37
post #24

Earlier quoted context omitted.

> This article has nothing to do with the current tête-à-tête with the Pentagon. The article yes, but we cannot be sure about its topic. We definitely cannot claim that they are unrelated. We don't know. It's possible that the two things have nothing to do with each other. It's also possible that they wanted to prevent worse requests and this was a preventive measure.

This is something they've been working on "in recent months". The Pentagon thing was today . This cannot have been caused by that, unless they've also invented time travel.

You heard about the Pentagon thing today. Doesn't mean it wasn't started because of political pressure.

Re: Anthropic drops flagship safety pledge

#38

This headline unfortunately offers more smoke than light. This article has nothing to do with the current tête-à-tête with the Pentagon. It is discussing one specific change to Anthropic's "Responsible Scaling Policy" that the company publicly released today as version "3.0".

I consider this a bigger deal than the Pentagon thing.

While not surprising at the least, it still kind of crazy that literal pdf files in charge is not concerning, but this is.

I just hope something happens to USA before it can do damage to the world.

Post reply on HN