Live data from Hacker News

Anthropic drops flagship safety pledge

time.com

391–400 of 716 posts

Re: Anthropic drops flagship safety pledge

#392
post #351

Worth checking this post from someone who actually has worked on this change: > I take significant responsibility for this change. https://www.lesswrong.com/posts/HzKuzrKfaDJvQqmjh/responsibl...

This guy from Effective Altruism pivoted away from helping the poor to help try to control AI from being a terminator type entity and then pivoted to being, ah, its okay for it to be a terminator type entity. > Holden Karnofsky, who co-founded the EA charity evaluator GiveWell, says that while he used to work on trying to help the poor, he switched to working on artificial intelligence because of the “stakes”: > “The…

> then pivoted to being, ah, its okay for it to be a terminator type entity.

Isn’t that the opposite of what he’s saying? He’s saying it could become that powerful, and given that possibility it’s incredibly important that we do whatever we can to gain more control of that scenario

Re: Anthropic drops flagship safety pledge

#393

Public benefit corporations in the AI space have become a farce at this point. They're just regular corporations wearing a different hat, driven by the same money dynamics as any other corp. They have no ability to balance their stated "mission" with their drive for profit. When being "evil" is profitable and not-evil is not, guess which road they'll take...

Well, now I'm wondering, if the company was chartered with the public benefit in mind, could you not sue if they don't follow through with working in the public interest? If regular corporations are sued for not acting in the interests of shareholders, that would suggest that one could file a suit for this sort of corporate behavior. I'm not even a lawyer (I don't even play one on TV) and public benefit corporations…

I really don’t see it. PBCs are dual purpose entities - under charter, they have a dual purpose of making profit while adding some benefit to society. Profit is easy to define; benefit to society is a lot more difficult to define. That difficulty is reflected at the penalty stage where few jurisdictions have any sort of examination of PBC status.

This is what we were all going on about 15 years ago when Maryland was the first state to make PBCs legal. We got called negative at the time.

Re: Anthropic drops flagship safety pledge

#394

Earlier quoted context omitted.

In general public benefit corporations and non-profits should have a very modest salary cap for everybody involved and specific public-benefit legally binding mission statements. Anybody involved should also be prohibited from starting a private company using their IP and catering to the same domain for 5-10 years after they leave. Non-profits where the CEO makes millions or billions are a joke. And if e.g. your miss…

It’s not the CEO’s fault - they had to take all that money to keep their org a non-profit. B corps are like recycling programs, a nice logo.

Don't they get tax breaks and more lax operating requirements? I don't think this is just an image thing.

Re: Anthropic drops flagship safety pledge

#395
post #351

Worth checking this post from someone who actually has worked on this change: > I take significant responsibility for this change. https://www.lesswrong.com/posts/HzKuzrKfaDJvQqmjh/responsibl...

I genuinely believe that website is responsible for a lot of the worst ideas currently permeating the technology sector.

Re: Anthropic drops flagship safety pledge

#396
post #372

Earlier quoted context omitted.

Well before Anthropic thought they were God's gift to AI; the chosen ones protecting humanity. With the latest competing models they are now realizing they are an "also" provider. Sobering up fast with ice bucket of 5.3-codex, Copilot, and OpenCode dumped on their head.

Hello sama

Sama-sama.

Re: Anthropic drops flagship safety pledge

#397
post #361

[flagged]

I hate comments anthropomorphizing LLMs. You are just asking a token producing system to produce tokens in a way that optimises for plausibility. Whatever it writes has no relation to its inner workings or truths. It doesn't "believe". It has no "intent". It cannot "admit". Steering a LLM to say anything you want is the defining characteristic of an LLM. That's how we got them to mimic chatbots. It's not clear there…

“believe” yes in the sense that my program believes x=7. Actually when it goes to read it maybe the bit flipped. Everything on machines is probabilistic that’s a tautology. However we have windowed bounds on valid output, and Claude being able to build a context in which its next decisions are trained on it being an angry vengeful god is not inside that window. That’s what “safe” means, as one of many possible examples.

Inner workings were determined by me, not the LLM. It assisted in generating inputs which had 100% boolean results in the output.

Re: Anthropic drops flagship safety pledge

#398
post #364
post #335

Hopefully this is the short-term move made only under duress so that they can file a lawsuit.

the article specifically says: > The policy change is separate and unrelated to Anthropic’s discussions with the Pentagon, according to a source familiar with the matter.

I'm not fond of this trend of stating a position and attributing it to "a source familiar with the situation"

It combines interpretation of meaning with ambiguity to allow the reporter to assert anything they want. The ambiguity is there to protect the identity of the source but it has to be a more discrete disclosure of information in return. If you can't check the person you can still check what they said.

I would be ok with direct quotes from an anonymous source. That removes the interpretation of meaning at least.

As it is written, it would not be inaccurate to say this if their source was the lesswrong post, or even an earlier thread here on HN.

Phrasing "A source with direct knowledge of the situation" might remove some of the leeway for editorialising, but without sharing what the source actually said, it opens the door to saying anything at all and declaring "That's what I thought they meant" when challenged.

It's unfalsifyible journalism.

Re: Anthropic drops flagship safety pledge

#399
post #351

Worth checking this post from someone who actually has worked on this change: > I take significant responsibility for this change. https://www.lesswrong.com/posts/HzKuzrKfaDJvQqmjh/responsibl...

I genuinely believe that website is responsible for a lot of the worst ideas currently permeating the technology sector.

pretty much the intellectual equivalent of looksmaxxing
Post reply on HN