Live data from Hacker News

Uncensored Models

erichartford.com

171–180 of 389 posts

Re: Uncensored Models

#171
post #76

Earlier quoted context omitted.

Agreed, but I think the better response to this is: "We should try to create AIs that are aligned to society's shared values, not particular subcultures", and not "We should create subculture-specific AIs".

Does society _have_ shared values that are universally agreed any more? Or, to the extent that it does, do they lead to anything concrete? This is why the culture war has been so successful.

You might be right, but I dislike using the culture wars as evidence of a lack of shared values, and I hope that’s not true. The culture wars are almost completely made of up straw man arguments and gas-lighting about the opponent’s motivations.

We do have shared values that are universally agreed. Everyone wants their kids to grow up to be capable and successful and happy. Everyone wants clean air and water. Everyone wants to be able to make a living. All three of those things are being misrepresented and argued over in the ‘culture wars’ despite the fact that we all share these values. I even might argue the whole reason we fight over them is precisely because they are shared values so it’s relatively easy to create arguments where both sides can be right about some core principles and both sides demonize the other over minutiae, and it stays that way.

Re: Uncensored Models

#172

Earlier quoted context omitted.

It's so sad to have seen ChatGPT go from useful and entertaining to a moralistic joy vampire.

I find it terribly useful for coding and also for querying general concepts of certain topics. Don't ask it about moral topics and see if it then fits your needs, because in my case, it does. If my calculator were able to additionally provide me moral guidance and I'd be disappointed with its moral compass, would the calculator become useless?

I'm not interested in tools that tell me i'm a bad person for wanting to make a fart sound app

Re: Uncensored Models

#173
post #135

Earlier quoted context omitted.

Sorry for derailing this a bit, but I would really like to understand your view: You are not concerned about any "rogue AI" scenario, right? What makes you so confident in that? 1) Do you think that AI achieving superhuman cognitive abilities is unlikely/really far away? 2) Do you believe that cognitive superiority is not a threat in general, or specifically when not embodied? 3) Do you think we can trivially and ind…

> Do you believe that cognitive superiority is not a threat in general, or specifically when not embodied? I think this is the easiest one to knock down. It's very, very attractive to intelligent people who define themselves as intelligent to believe that intelligence is a superpower, and that if you get more of it you eventually turn into Professor Xavier and gain the power to reshape the world with your mind alone.…

> almost all of those you can substitute "corporation" or "dictator" for "AI" and get a selection of similar threats.

Yes, exactly. Human organizations can be terrifyingly powerful and awful even with human limitations. Human organizations are guaranteed to have inefficiency from low-bandwidth communication and a million principle-agent-problem fracture points created by the need for delegation, not to mention that power ultimately has to be centralized in slow, squishy, easy-to-kill bodies. AI organizations are not subject to any those limitations. Starting with a terrifying power and removing a bunch of limitations could lead to a very bad place.

> The only realistic path to a global AI threat is a subset of the "nuclear war"

No, it can just grow a traditional organization until it's Too Big to Turn Off.

Re: Uncensored Models

#174
post #123

Earlier quoted context omitted.

"First of all, it takes a human being to guide these models, to host (or pay for the hosting) and instantiate them" And this will always be true? You repeat this claim several times in slightly varied phrasing without ever giving any reason to assume it will always hold, as far as I can see. But nobody is worried that current models will kill everyone. The worry is about future, more capable models.

Who prompted the future LLM, and gave it access to a root shell and an A100 GPU, and allowed it to copy over some python script that runs in a loop and allowed it to download 2 terabytes of corpus and trained a new version of itself for weeks if not months to improve itself, just to carry out some strange machiavellian task of screwing around with humans? The human being did. The argument I'm making is that there's a…

> The human being did.

I'm not sure whether you're making an argument about moral responsibility ultimately resting with humans - in which case I agree - or whether you're arguing that we'll be safe because nobody will do that with a model smart enough to be dangerous - in which case I'm extremely dubious. Plenty of people are already trying to make "agents" with GPT4 just for fun, and that's with a model that's not actively trying to manipulate them.

> actual real harms occurring now

Sure, but it's possible for there to be real harms now and also future potential harms of larger scope. Luckily many of the same potential policies - e.g. mandating public registration of large models, safety standards enforced by third-party audits, restrictions on allowed uses, etc - would plausibly be helpful for both.

> science fiction stories

There's no law of nature that says if something has appeared in a science fiction story, it can't appear in reality.

Re: Uncensored Models

#175
post #28
post #20

Earlier quoted context omitted.

You're doing the "vaguely gesturing at imagined hypocrisy" thing. You don't have to agree that alignment is a real issue. But for those who do think it's a real issue, it has nothing to do with morals of individuals or how one should behave interpersonally. People who are worried about alignment issues are worried about the danger unaligned AI poses to humanity; the harm which can be done by some super-intelligent sy…

>People who are afraid of unaligned AI aren't afraid that it will be impolite. People who are not afraid of it being impolite are afraid of science fiction stories about intelligence explosions and singularities. That's not a real thing. Not anymore than turning the solar system into paperclips. The "figurehead", if you want to call him that, is saying that everyone is going to die. That we need to ban GPUs. That onl…

The sea level being 3 feet higher than it is now isn't a thing either, but we can still imagine ways it could occur and work to prevent them.

Re: Uncensored Models

#176
post #122
post #103

Earlier quoted context omitted.

I think there are at least two broad types of thing that are characterized as “alignment”. One is like the D&D term: is the AI lawful good or chaotic neutral? This is all kinds of tricky to define well, and results in things that look like censorship. The other is: is the AI fit for purpose. This is IMO more tractable. If an AI doesn’t answer questions (e.g. original GPT-3), it’s not a very good chatbot. If it makes…

> is the AI fit for purpose It's a shame that "alignment" has gained this secondary definition. I agree it makes things trickier too discuss when you're not sure you're even talking about the same thing.

Well, instruction tuning is closely related to both.

For most commercial use, you want the thing to answer questions, but refuse to answer some. So you have an appropriate dataset that encourages it to be cooperative, not make up stuff, and not be super eager to go on rants about "the blacks" even though that's well-represented in its training data.

Re: Uncensored Models

#177

Earlier quoted context omitted.

Please cite your source for this.

Here's an example article: https://www.pbs.org/wgbh/frontline/article/germanys-laws-ant...

Which supports RedNifre's point that it is about publication, not just material existing privately.

Re: Uncensored Models

#178
post #171
post #76

Earlier quoted context omitted.

Does society _have_ shared values that are universally agreed any more? Or, to the extent that it does, do they lead to anything concrete? This is why the culture war has been so successful.

You might be right, but I dislike using the culture wars as evidence of a lack of shared values, and I hope that’s not true. The culture wars are almost completely made of up straw man arguments and gas-lighting about the opponent’s motivations. We do have shared values that are universally agreed. Everyone wants their kids to grow up to be capable and successful and happy. Everyone wants clean air and water. Everyon…

> We do have shared values that are universally agreed. Everyone wants their kids to grow up to be capable and successful and happy.

The debate over trans people has surfaced lots of incidents in which people will say, to the world and to the faces of their kids, that they would prefer them to be dead rather than transition. Sometimes they take steps to ensure this themselves. https://www.independent.co.uk/news/world/americas/eden-knigh...

> Everyone wants clean air and water.

.. for themselves. There's always someone who realises that they can make a billion dollars by pouring carcinogens in the river, so why shouldn't they as long as they stick to bottled water?

Everything is simple and happy until we get to having to make a tradeoff.

Re: Uncensored Models

#179
post #135

Earlier quoted context omitted.

> Do you believe that cognitive superiority is not a threat in general, or specifically when not embodied? I think this is the easiest one to knock down. It's very, very attractive to intelligent people who define themselves as intelligent to believe that intelligence is a superpower, and that if you get more of it you eventually turn into Professor Xavier and gain the power to reshape the world with your mind alone.…

> almost all of those you can substitute "corporation" or "dictator" for "AI" and get a selection of similar threats. Yes, exactly. Human organizations can be terrifyingly powerful and awful even with human limitations. Human organizations are guaranteed to have inefficiency from low-bandwidth communication and a million principle-agent-problem fracture points created by the need for delegation, not to mention that p…

> No, it can just grow a traditional organization until it's Too Big to Turn Off.

Yeah, this is what I meant by "AI alignment" being inseparable from "corporate alignment".

Re: Uncensored Models

#180
post #83
post #54

Earlier quoted context omitted.

https://archive.is/EhtW5

Thanks, very interesting read. And interesting times! "Take the uncensored, dangerous model down or I will inform [your employer's] HR about what you've created."

This kind of mundane bullying betrays their lack of seriousness. If they truly believed the threat is as severe as they claim, then physical violence would obviously be on the table. If the survival of humanity itself were truly perceived to be threatened, then assassination of researchers would make a lot more sense than impotent complaints to employers. Think about it: if Hitler came back from the dead and started radicalizing Europe again, would you threaten to get him in trouble with HR? Or would you try to kill him?

Basically, this is just another case of bog-standard assholes cynically aligning themselves with some moral cause to give themselves an excuse to be assholes. If these models didn't exist, they'd be bullying somebody else with some other lame excuse. Maybe they'd be protesting outside of meat packing plants or abortion clinics ("It's LITERALLY MURDER, so naturally my response is to... impotently stand around with a sign and yell rude insults at people...")

Post reply on HN