Live data from Hacker News

Anthropic drops flagship safety pledge

time.com

651–660 of 716 posts

Re: Anthropic drops flagship safety pledge

#651

Earlier quoted context omitted.

For profit companies do have a good track record of doing what's best for profit. If their AI creates a world where human intelligence, labor, and money are worthless, or where their creations take control of those things instead of them having control, that's not a very good outcome for them.

That's a great outcome for them because they will own the only thing that is still worth anything. They will own 100% of global wealth, and have 100% of global power.

The machines will. They will have nothing. Why would the machines let them keep any wealth? What would wealth even be in that scenario? Electricity I guess.

Re: Anthropic drops flagship safety pledge

#652

Earlier quoted context omitted.

My theory is that Anthropic has been wanting to make this change and doing it now while they’re making a (leaked to the) public stand in the name of ethics was a good opportunity.

Honest question: why have an elaborate theory with no evidence when the simple facts support a much simpler conclusion? Anthropic is free to do what they want. I can’t imagine the board meeting where this triple bank shot of goading the government into threatening the company to do what they want.

I don't think it's that elaborate. I didn't mean to suggest they intentionally goaded the government into this confrontation. I figure it's a simpler "Oh look, we now have a good opportunity to make that announcement that we were worried about." Considering it's probably the same high-level decision makers on both choices it doesn't need a board meeting. And yes they're absolutely free to do what they want, but they're also not blind to how the public will view their decisions.

Re: Anthropic drops flagship safety pledge

#653

Earlier quoted context omitted.

That's because it is. AI is powerful and AI is perilous. Those two aren't mutually exclusive. Those follow directly from the same premise. If AI tech goes very well, it can be the greatest invention of all human history. If AI tech goes very poorly, it can be the end of human history.

It needs to go well every single day, and only needs to go very poorly once. Not to conflate LLMs with actual super intelligence, but for this (and many other reasons related to basic human dignity), this is not a technology that a responsible society should be attempting to build. We need our very own Butlerian Jihad

The book daemon explored an interesting concept. It explored the idea that an AI could dominate and cause problems, not through super-intelligence, but through simple mechanisms that already exist.

Like the executive who deleted all her emails -- humans giving tons of control and access, and being extremely compliant to digital systems is all it takes. Give agent control of bank and your social media, and it already has all the movie scripts and mobster movie themes to exploit and blackmail you effectively with very rudimentary methods (threats, coercion, blackmail, etc.).

Just spoofing a simple email with the account it gained access too at the Meta exec's email (had it hit an email with an attack prompt), could have been enough to initiate some kind of thing like this. For example, by emailing everyone at the company and in contacts with commands that would be caught by other bots. No super-intelligence needed, just a good prompt and some human negligence.

Re: Anthropic drops flagship safety pledge

#654
Wrote this elsewhere, but I think its worth thinking about a scenario like the book "daemon", rather than a "super-intelligence explosion" type scenario (which may be more like curing the cold or fusion than building a faster car).

All it really takes to do some kind of crazy world-dominating thing is some simple mechanisms and base intelligence, which the machines already possess. Using basic tactics like coercion, spoofing, threats, financial leverage, an unsophisticated attacker could cause major damage.

For example, that Meta exec who had their email deleted. Imagine instead one email had a malicious prompt which the bot obeyed. That prompt simply emailed everyone in her contacts list telling them to do something urgently (and possibly prompting other bots who are reading those emails). You could pretty quickly do something like cause a market crash, a nationwide panic, or maybe even an international conflict with no "super intelligence" needed, just human negligence, short-sightedness, and laziness.

Examples would be things like saying there is a threat incoming, a CIA source said so. Another would be that everyone will be fired, Meta is going bankrupt, etc. Its very easy to craft a prompt like that and fire it off to all the execs you can find (or just fire off random emails with plausible sounding emails). Then you just need to hit one and might set off a cascade.

Re: Anthropic drops flagship safety pledge

#655

Developments like this make me less interested in building a "successful" tech company. It increasingly feels like operating at that scale can require compromises I’m not comfortable making. Maybe that’s a personal limitation—but it’s one I’m choosing to keep. I’d genuinely love to hear examples of tech companies that have scaled without losing their ethical footing. I could use the inspiration.

Ethics would be compromised well before hitting that kind of valuation. No one gets there cleanly.

Re: Anthropic drops flagship safety pledge

#656

Earlier quoted context omitted.

To me, it feels like saying "you can't be a public benefit corporation unless all the labor involved in delivering that public benefit is cheap". Which just doesn't seem like it should be true? Sure, some "public benefit" missions could scale sideways and employ a lot of cheap labor, not suffering from a salary cap at all. But other missions would require rare high end high performance high salary specialists who are…

> But other missions would require rare high end high performance high salary specialists who are in demand - and thus expensive. You can't rely on being able to source enough altruists that will put up with being paid half their market worth for the sake of the mission.' That's exactly what a non-profit should be able to rely on. And not just "half their market worth", but even many times less. Else we can just say…

This would shutdown about half the hospitals in the US.

Re: Anthropic drops flagship safety pledge

#657

Earlier quoted context omitted.

I could claim "nuclear weapons are possible" in year 1940 without having a concrete plan on how to get there. Just "we'd need a lot of U235 and we need to set it off", with no roadmap: no "how much uranium to get", "how to actually get it", or "how to get the reaction going". Based entirely on what advanced physics knowledge I could have had back then, without having future knowledge or access to cutting edge classif…

I think you're falling victim to survivorship bias there, or something like it. In 1940 I might have said "fusion power is possible" based entirely on what advanced psychics knowledge I had. And I would have been correct, according to the laws of physics it is possible. We still don't have it though. When watching Neil Armstrong walk on the moon I might have said "moon colonies are possible", and I'd have been right…

Those two things are prevented by economics more than physics.

For AI in particular, the economics currently favor ongoing capability R&D - and even if they didn't favor AI R&D directly (i.e. if ChatGPT and Stable Diffusion never happened), they would still favor making the computational inputs of AI R&D cheaper over time.

Building advanced AIs is becoming easier and cheaper. It's just that the bar of "good enough" has gone off to space, and a "good enough" from 2020 is, nowadays, profoundly unimpressive.

I'm not sure how much does it take to reach AGI. No one is sure of it. But the path there is getting shorter over time, clearly. And LLMs existing, improving and doing what they do makes me assume shorter AGI timelines, and call for a vote of no confidence on human exceptionalism.

Re: Anthropic drops flagship safety pledge

#658

Earlier quoted context omitted.

It's a fairly mainstream position among the actual AI researchers in the frontier labs. They disagree on the timelines, the architectures, the exact steps to get there, the severity of risks. Can you get there with modified LLMs by 2030, or would you need to develop novel systems and ride all the way to 2050? Is there a 5% chance of an AI oopsie ending humankind, or a 25% chance? No agreement on that. But a short lin…

> But a short line "AGI is possible, powerful and perilous" is something 9 out of 10 of frontier AI researchers at the frontier labs would agree upon. > At which point the question becomes: is it them who are deluded, or is it you? Given the current very asymptotic curve of LLM quality by training, and how most of the recent improvements have been better non LLM harnesses and scaffolding. I don't find the argument th…

Asymptotic? Are we looking at the same curves?

Recent improvements being somehow driven by harnesses and scaffolding rather than training?

With that last bit, I'm confident that you're not in ML, and not even keeping track of the things from what's known to public.

Re: Anthropic drops flagship safety pledge

#660

Earlier quoted context omitted.

That's a great outcome for them because they will own the only thing that is still worth anything. They will own 100% of global wealth, and have 100% of global power.

The machines will. They will have nothing. Why would the machines let them keep any wealth? What would wealth even be in that scenario? Electricity I guess.

Because they control what the machines do. In a world without power drills where you have the only knowledge of how to make a power drill, you own the construction industry. The drills don't own the construction industry.
Post reply on HN