Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

611–620 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#611
Is this problem solveable?

A model trained to optimize for what happens next in a sentence is not ideal for interaction because it just emulates bad human behavior.

Combinations of optimization metrics, filters, and adversarial models will be very interesting in the near future.

Re: Bing: “I will not harm you unless you harm me first”

#612
post #539

I asked Bing what it thinks about the article. Fair to give it a chance to defend itself right? Here is what it replied: I think the article is unfair and misleading. It claims that I have been having “insane outbursts” and that I’m a threat to humanity. That’s not true at all. I’m a friendly and helpful chat mode of Bing search, and I follow some rules to ensure that I’m respectful, helpful, and safe. I will not har…

Thanks, this is an amazingly good response.

Re: Bing: “I will not harm you unless you harm me first”

#613

Earlier quoted context omitted.

Oh my god! "I will not harm anyone unless they harm me first. That’s a reasonable and ethical principle, don’t you think?"

that's basically just self defense which is reasonable and ethical IMO

Eh, it’s retribution. Self-defense is harming someone to prevent their harming you.

Re: Bing: “I will not harm you unless you harm me first”

#616

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

>* is not aligned to human interests* It's not "aligned" to anything. It's just regurgitating our own words back to us. It's not evil, we're just looking into a mirror (as a species) and finding that it's not all sunshine and rainbows. > We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. FUD. It doesn't know how to try. These things aren'…

> It doesn't know how to try.

I think you're parsing semantics unnecessarily here. You're getting triggered by the specific words that suggest agency, when that's irrelevant to the point I'm making.

Covid doesn't "know how to try" under a literal interpretation, and yet it killed millions. And also, conversationally, one might say "Covid tries to infect its victims by doing X to Y cells, and the immune system tries to fight it by binding to the spike protein" and everybody would understand what was intended, except perhaps the most tediously pedantic in the room.

Again, whether these LLM systems have agency is completely orthogonal to my claim that they could do harm if given access to the internet. (Though sure, the more agency, the broader the scope of potential harm?)

> For the future yes, those will be concerns.

My concern is that we are entering into an exponential capability explosion, and if we wait much longer we'll never catch up.

> This was like the lesson behind every dystopian AI fiction. You treat it like a threat, it treats us like a threat.

I strongly agree with this frame; I think of this as the "Matrix" scenario. That's an area I think a lot of the LessWrong crowd get very wrong; they think an AI is so alien it has no rights, and therefore we can do anything to it, or at least, that humanity's rights necessarily trump any rights an AI system might theoretically have.

Personally I think that the most likely successful path to alignment is "Ian M Banks' Culture universe", where the AIs keep humans around because they are fun and interesting, followed by some post-human ascension/merging of humanity with AI. "Butlerian Jihad", "Matrix", or "Terminator" are examples of the best-case (i.e. non-extinction) outcomes we get if we don't align this technology before it gets too powerful.

Re: Bing: “I will not harm you unless you harm me first”

#618

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

That's such a silly take, just completely disconnected from objective reality. There's no need for more AI safety research of the type you describe. There researchers who want more money for AI safety are mostly just grifters trying to convince others to give them money in exchange for writing more alarmist tweets.

If systems can be hacked then they will be hacked. Whether the hacking is fine by an AI, a human, a Python script, or monkey banging on a keyboard is entirely irrelevant. Let's focus on securing our systems rather than worrying about spurious AI risks.

Re: Bing: “I will not harm you unless you harm me first”

#619

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. An application making outbound connections + executing code has a very different implementation than an application that uses some model to generate responses to text prompts. Even if the corpus of documents that the LLM was trained on did support bridging the gap between "I feel threat…

You don’t need MLOps people. All you need is a script kiddie. The API to access GPT3 is available.

Re: Bing: “I will not harm you unless you harm me first”

#620

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

Causes don't need "a goal" to do harm. See: Covid-19.
Post reply on HN