Live data from Hacker News

Anthropic drops flagship safety pledge

time.com

521–530 of 716 posts

Re: Anthropic drops flagship safety pledge

#522

Anthropic's CEO Dario has annoyed me to no end with his "AI will take all the jobs in 6 months" doomer speeches on every podcast he graces his presence with.

He’s an e/acc guy. That should tell you everything. And maybe the incredibly awkward behavior and demeanor.

Re: Anthropic drops flagship safety pledge

#523

I was wondering if it was because of heavy-handedness of the administration, but apparently: > The policy change is separate and unrelated to Anthropic’s discussions with the Pentagon, according to a source familiar with the matter. Their core argument is that if we have guardrails that others don't, they would be left behind in controlling the technology, and they are the "responsible ones." I honestly can't compreh…

Let's suppose I believe them, that's still a bad idea.

The reason Claude became popular is because it made shit up less often than other models, and was better at saying "I can't answer that question." The guardrails are quality control.

I would rather have more reliable models than more powerful models that screw up all the time.

Re: Anthropic drops flagship safety pledge

#524

Earlier quoted context omitted.

It's a fairly mainstream position among the actual AI researchers in the frontier labs. They disagree on the timelines, the architectures, the exact steps to get there, the severity of risks. Can you get there with modified LLMs by 2030, or would you need to develop novel systems and ride all the way to 2050? Is there a 5% chance of an AI oopsie ending humankind, or a 25% chance? No agreement on that. But a short lin…

Sure, when you get rid of the timelines and the methods we'll use to get there, everyone agrees on everything. But at that point it means nothing. Yeah, AGI is possible (say the people who earn a salary based on that being true). Curing all known diseases is possible too. How will we do that? Oh, I don't know. But it's a thing that could possibly happen at some point. Give me some investment cash to do it. If you cla…

I could claim "nuclear weapons are possible" in year 1940 without having a concrete plan on how to get there. Just "we'd need a lot of U235 and we need to set it off", with no roadmap: no "how much uranium to get", "how to actually get it", or "how to get the reaction going". Based entirely on what advanced physics knowledge I could have had back then, without having future knowledge or access to cutting edge classified research.

Would not having a complete foolproof step by step plan to obtaining a nuclear bomb somehow make me wrong then?

The so-called "plan" is simply "fund the R&D, and one of the R&D teams will eventually figure it out, and if not, then, at least some of the resources we poured into it would be reusable elsewhere". Because LLMs are already quite useful - and there's no pathway to getting or utilizing AGI that doesn't involve a lot of compute to throw at the problem.

Re: Anthropic drops flagship safety pledge

#525
post #351

Worth checking this post from someone who actually has worked on this change: > I take significant responsibility for this change. https://www.lesswrong.com/posts/HzKuzrKfaDJvQqmjh/responsibl...

This guy from Effective Altruism pivoted away from helping the poor to help try to control AI from being a terminator type entity and then pivoted to being, ah, its okay for it to be a terminator type entity. > Holden Karnofsky, who co-founded the EA charity evaluator GiveWell, says that while he used to work on trying to help the poor, he switched to working on artificial intelligence because of the “stakes”: > “The…

Getting SBF vibes from this. "Earn to give" is an inherently flawed philosophy.

Re: Anthropic drops flagship safety pledge

#526

Earlier quoted context omitted.

That's because it is. AI is powerful and AI is perilous. Those two aren't mutually exclusive. Those follow directly from the same premise. If AI tech goes very well, it can be the greatest invention of all human history. If AI tech goes very poorly, it can be the end of human history.

Same with everything, right? You could say the same with nukes, electricity, internet, the computer, etc... But if you look at it without paying attention to the "ultimate tool for humanity" hype, it doesn't really look that much of a threat or a salvation. It won't end civilization for dropping the guardrails, but it will surely enable bad actors to do more damage than before (mass scams, blackmail, deepfake nudes,…

One difference is the very real possibility that AI will not just be a "tool for humanity", but a collection of actors with real power and goals. Robert Miles has an approachable explanation here: https://www.youtube.com/watch?v=zATXsGm_xJo

Re: Anthropic drops flagship safety pledge

#527

This drama arc of “I used to be so pure and good, but others made me evil” is so tiring. I really miss the nerd profile who cared a lot more about tech and science, and a lot less about signaling their righteousness. How did we get so religious/narcissistic so quickly and as a whole?

One might argue that this corresponds to the general shift of the political left towards these things. Old pre-turn-of-century tech was a much more libertarian left. Notice how a lot of the 50-something gen-X CEOs (and others) were once "left" but are now hated by that group, and more likely to go over to Trumpism. Obvious case in point: Elon

The entire playing field is kinda dissapointing, left or right. Which do you wanna be, self-righteous preening snob or batshit macho man?

I'm going for a blend, myself

Re: Anthropic drops flagship safety pledge

#528
I'm still a little fuzzy on what "safety" even means anymore. If someone could explain it, that would be great.

Because at this point, it's too broad to be defined in the context of an LLM, so it feels like they removed a blanket statement of "we will not let you do bad things" (or "don't be evil"), which doesn't really translate into anything specific.

Re: Anthropic drops flagship safety pledge

#529

Earlier quoted context omitted.

John Good's quote is pretty myopic, it assumes machines make better machines based on being "ultraintelligent" instead of learning from environment-action-outcome loop. It's the difference between "compute is all you need" and "compute+explorative feedback" is all you need. As if science and engineering comes from genius brains not from careful experiments.

> As if science and engineering comes from genius brains not from careful experiments 100% this. How long were humans around before the industrial revolution? Quite a while

Science and engineering didn't begin with the Industrial Revolution. See: https://en.wikipedia.org/wiki/Great_Pyramid_of_Giza

Re: Anthropic drops flagship safety pledge

#530

Earlier quoted context omitted.

That's because it is. AI is powerful and AI is perilous. Those two aren't mutually exclusive. Those follow directly from the same premise. If AI tech goes very well, it can be the greatest invention of all human history. If AI tech goes very poorly, it can be the end of human history.

> If AI tech goes very well The IF here is doing some very heavy lifting. Last I checked, for profit companies don't have a good track record of doing what's best for humanity.

For profit companies do have a good track record of doing what's best for profit. If their AI creates a world where human intelligence, labor, and money are worthless, or where their creations take control of those things instead of them having control, that's not a very good outcome for them.
Post reply on HN