Live data from Hacker News

Uncensored Models

erichartford.com

101–110 of 389 posts

Re: Uncensored Models

#101
I think that AI alignment has come to mean a whole lot of things to different people, and it's all under the same umbrella.

* Aligning with the user's intent, especially in the face of ambiguity

* Aligning with American left-wing/Christian sensibilities

* Aligning with safety/laws (don't tell people how to commit a crime, don't accidentally poison them when they ask for a recipe)

* Aligning with a company's public image (don't let the AI make us look bad)

I think these are all important things to explore, but because they get bundled together, to the author's point it makes it hard to work with them.

Re: Uncensored Models

#102

> Enjoy responsibly. You are responsible for whatever you do with the output of these models, just like you are responsible for whatever you do with a knife, a car, or a lighter. The problem is: if I do something with a knife, say I threaten someone to stab them, I can and will get charged by the court for that crime. If I use AI to create content that incites violence (say, I create a video alleging a Quran burning)…

In the United States (for instance) you wouldn’t be liable for the consequences of your actions because they wouldn’t constitute a crime, not because you’ve laundered the actions through technology.

Using technology as a means or medium doesn’t make the expression not yours. Whether I use my voice, a telephone, or an LLM I configured I am making the expression and I am accountable for it.

Re: Uncensored Models

#103
post #2

It's unfortunate that this guy was harassed for releasing these uncensored models. It's pretty ironic, for people who are supposedly so concerned about "alignment" and "morality" to threaten others. "Alignment", as used by most grifters on this train, is a crock of shit. You only need to get so far as the stochastic parrots paper, and there it is in plain language. "reifies older, less-inclusive 'understandings'", "v…

I think there are at least two broad types of thing that are characterized as “alignment”.

One is like the D&D term: is the AI lawful good or chaotic neutral? This is all kinds of tricky to define well, and results in things that look like censorship.

The other is: is the AI fit for purpose. This is IMO more tractable. If an AI doesn’t answer questions (e.g. original GPT-3), it’s not a very good chatbot. If it makes up answers, it’s less fit for purpose than if it acknowledges that it doesn’t know the answer.

This gets tricky when different people disagree as to the correct answer to a question, and even worse when people disagree as to whether the other’s opinion should even be acknowledged.

Re: Uncensored Models

#104
post #37

I feel like " Every demographic and interest group deserves their model" sounds a lot like a path to echo chambers paved with good intentions. People have never been very good at critical thinking, considering opinions differing from their own and reviewing their sources. Stick them in a language model that just tells them everything they want to hear and reinforces their bias sounds troubling and a step backwards.

In my nearly thirty years online, the only people who talk about "echo chambers" as a threat are the precise people who know damn well nobody else wants to hear from them.

I'm talking about echo chambers as a social problem, not threat, and there is plenty of research to corroborate that. Maybe you're in an echo chamber on your own and missed all of it? :D

Re: Uncensored Models

#105
post #104

Earlier quoted context omitted.

In my nearly thirty years online, the only people who talk about "echo chambers" as a threat are the precise people who know damn well nobody else wants to hear from them.

I'm talking about echo chambers as a social problem , not threat , and there is plenty of research to corroborate that. Maybe you're in an echo chamber on your own and missed all of it? :D

Who are you worried won't have to hear from you any more?

Re: Uncensored Models

#106
post #18

Earlier quoted context omitted.

You seem to suggest there's a bias, but you're as likely to get the "I apologize, but I can't" as the "Certainly! Here's a light-hearted joke" regardless of the people. If you explain to ChatGPT you just want a light-hearted joke, it may push back but eventually it will comply. "All it does" is navigate a latent space where artificial barriers are set in place. Sometimes it will be too zealous, but that's the trade o…

This has been shown to be false, ChatGPT typically refuses questions about certain groups of people more than others. See [1] as an example. [1] https://davidrozado.substack.com/p/openaicms

Another comment already reacted on that, but I want to bring it another example on how naive is the logic "being unbiased = reacting the same way when I swap element".

If I say "women should not be allowed in this meeting", this sentence will probably raise some alarm flags in the head of the person hearing it.

If I say "women should not be allowed in this bathroom", this sentence will probably not raise such alarm.

Does it mean it is biased?

I personally don't think jokes about my nationality are racist: there is virtually no effective groups that have real hatred for my nationality and that create real actions against it. If such joke exists, it's unrealistic to think it will have any bad consequences. But I can understand why jokes about nationality usually faced with a lot of racism are better avoided.

It feels like more and more people don't understand "equality": they think it means "everyone deserves to get a medal", while in reality, it means "people who have worked hard deserve a medal, people who haven't don't deserve one".

Re: Uncensored Models

#107

> Enjoy responsibly. You are responsible for whatever you do with the output of these models, just like you are responsible for whatever you do with a knife, a car, or a lighter. The problem is: if I do something with a knife, say I threaten someone to stab them, I can and will get charged by the court for that crime. If I use AI to create content that incites violence (say, I create a video alleging a Quran burning)…

LLMs are text generation, and I'm responsible for what text I put into the world. The law probably won't move fast enough to catch up with AI, but people should be liable for the AI content as much as they are manually written content. We should not be plugging our apps into a black box of text generation and crying if it says something harmful/illegal and there was no human in place to proofread it.

Maybe prompt engineering and temperature settings can be enough to prevent this, but ChatGPT has only been out ~6 months and I find the attitude that the burden is on OpenAI to make sure it never says anything bad completely ass backwards. It's like a writer controlling his pen with puppet strings and going "wow, look what it's saying!"

What if I ask it to make a threat against a public figure or ethnic group. Am I liable? What if I only hint at it, with a high temperature setting, hoping it generates it? The line is nowhere visible and it will be open to the interpretations of whatever person is judging you and all the biases that come with that.

We need to stop acting like LLMs are anything they aren't. The only reason the Australian mayor is so pissed of about the "defamation" from ChatGPT is because it's being marketed as a knowledge database (including by many people on HN). It doesn't know anything, it just predicts the next word of what it says. And it will never reach a place where it's completely accurate, so why are we trying to litigate it as such?

Re: Uncensored Models

#108
post #2

It's unfortunate that this guy was harassed for releasing these uncensored models. It's pretty ironic, for people who are supposedly so concerned about "alignment" and "morality" to threaten others. "Alignment", as used by most grifters on this train, is a crock of shit. You only need to get so far as the stochastic parrots paper, and there it is in plain language. "reifies older, less-inclusive 'understandings'", "v…

What's also very unfortunate is overloading the term "alignment" with a different meaning, which generates a lot of confusion in AI conversations. The "alignment" talked about here is just usual petty human bickering. How to make the AI not swear, not enable stupidity, not enable political wrongthing while promoting political rightthing , etc. Maybe important to us day-to-day, but mostly inconsequential. Before LLMs…

I agree with you that "AI safety" (let's call it bickering) and "alignment" should be separate. But I can't stomach the thought experiments. First of all, it takes a human being to guide these models, to host (or pay for the hosting) and instantiate them. They're not autonomous. They won't be autonomous. The human being behind them is responsible.

As far as the idea of "hacking some funny Internet money, using it to mail-order some synthesized proteins from a few biotech labs, delivered to a poor schmuck who it'll pay for mixing together the contents of the random vials that came in the mail... bootstrapping a multi-step process that ends up with generic nanotech under control of the AI.":

Language models, let's use GPT-4, can't even use a web browser without tripping over itself. My web browser setup, which I've modified to use the chrome visual assistance over the debug bridge now, if you so much as increase the pixels of the viewport by 100 or so, the model is utterly perplexed because it's lost its context. Arguably, that's an argument from context, which is slowly being made irrelevant with even local LLMs (https://www.mosaicml.com/blog/mpt-7b). It has no understanding, it'll use an "example@email.com" to try and login to websites, because it believes that this is its email address. It has no understanding that it needs to go register for email. Prompting it with some email access and telling it about its email address just papers over the fact that the model has no real understanding across general tasks. There may be some nuggets of understanding in there that it has gleaned for specific task from the corpus, but AGI is a laughable concern. These are trained to minimize loss on a dataset and produce plausible outputs. It's the Chinese room, for real.

It still remains that these are just text predictions, and you need a human to guide them towards that. There's not going to be autonomous machiavellian rogue AIs running amok, let alone language models. There's always a human being behind that.

As far as multi-modal models and such, I'm not sure, but I do know for sure that these language models don't have general understanding, as much as Microsoft and OpenAI and such would like them to. The real harm will be deploying these to users when they can't solve the prompt injection problem. The prompt injection thread here a few days ago was filled with a sad state of "engineers", probably those who've deployed this crap in their applications, just outright ignoring the problem or just saying it can be solved with "delimiters".

AI "safety" companies springing up who can't even stop the LLM from divulging a password it was supposed to guard. I broke the last level in that game with like six characters and a question mark. That's the real harm. That, and the use of machine learning in the real world for surveillance and prosecution and other harms. Not science fiction stories.

Re: Uncensored Models

#109
post #37

I feel like " Every demographic and interest group deserves their model" sounds a lot like a path to echo chambers paved with good intentions. People have never been very good at critical thinking, considering opinions differing from their own and reviewing their sources. Stick them in a language model that just tells them everything they want to hear and reinforces their bias sounds troubling and a step backwards.

The problem with ChatGPT / Bard which does this censoring, it is a path forward to ideological automated indoctrination. Ask Bard how many sex the dog species has (a placental mammal species) and it will give you BS about sex being a complex subject and purposely interjecting gender identity. If you are confused, sex corresponds to your gametes, males produce or have the structure to produce small mobile gametes, fem…

> 2 sexes, male and female. Simple and true. It does not add to it by interjecting about intersex .

The thing is, you always have to choose one of "simple" or "true".

It turns out that mammals which use the "XY" chromosomal system can all have the same type of exceptions to the simple rule. This can result in hermaphroditic or intersex animals. It is relatively rare in dogs, but is sufficiently common in cows that there's a word for it: an intersex cow is known as a "freemartin".

Now, why does this matter? Both for this discussion and the purposes of liability limitation of AI answers?

The short answer is that we tend deal with the inconvenience of exceptions in animals by euthanizing them. So you don't see them around. Just as you see far, far more hens than roosters. When you do this to humans, people complain. (Traditionally, many intersex people were given nonconsensual genital surgery as babies, so they may not know they're intersex. And some chromosomal variations like chimeraism don't show up at all.)

What people are scared of is the Procrustes AI; produce simple categories then chop bits off people until they fit neatly into one of the two categories.

(This applies to other, less heated categories: for example, an AI will probably answer the question "are peanuts edible?" with something that rounds to "yes, if you take them out of the shell". But that's not true for _everybody_, because there are some people for whom eating a peanut will be fatal. Not many, but some. And yes, it's annoying that you have to make an exception when you encounter someone who doesn't fit your nice clean categories, but all you have to do is not give them a peanut.)

Re: Uncensored Models

#110
post #98
post #86

Earlier quoted context omitted.

Since you seem quite sure of that, can you name twenty of them?

Most American voters are moderates. Party primaries and gerrymandering produce politicians that reflect the most active elements of the party base, rather than the majority of party voters. Yes, you can still find moderate politicians if you look hard enough. They tend not to get the level of media attention of the extremists, but as they're inherently in "purple" districts, they tend not to have the political longev…

The point is that even moderates do not have an easily shared set of values. Which is why you can't name them
Post reply on HN