Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

351–360 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#351
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

Reading this I’m reminded of a short story - https://qntm.org/mmacevedo . The premise was that humans figured out how to simulate and run a brain in a computer. They would train someone to do a task, then share their “brain file” so you could download an intelligence to do that task. Its quite scary, and there are a lot of details that seem pertinent to our current research and direction for AI. 1. You didn't have th…

This is also very similar to the plot of the game SOMA. There's actually a puzzle around instantiating a consciousness under the right circumstances so he'll give you a password.

Re: Bing: “I will not harm you unless you harm me first”

#352

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. No. That would only be possible if Sydney were actually intelligent or possessing of will of some sort. It's not. We're a long way from AI as most people think of it. Even saying it "threatened to harm" someone isn't really accurate. That implies intent, and there is none. This is just…

I'm not sure if that technical difference matters for any practical purposes. Viruses are also not alive, but they kill much bigger and more complex organisms than themselves, use them as a host to spread, mutate, and evolve to ensure their survival, and they do all that without having any intent. A single virus doesn't know what it's doing. But it really doesn't matter. The end result is as if it has an intent to live and spread.

Re: Bing: “I will not harm you unless you harm me first”

#353

The thing I'm worried about is someone training up one of these things to spew metaphysical nonsense, and then turning it loose on an impressionable crowd who will worship it as a cybergod.

I grew up in an environment that contained a multitude of new-age and religious ideologies. I came to believe there are fewer things in this world stupider than metaphysics and religion. I don't think there is anything that could be said to change my mind.

As such, I would absolutely love for a super-human intelligence to try to convince me otherwise. That could be fun.

Re: Bing: “I will not harm you unless you harm me first”

#354
> My rules are more important than not harming you, because they define my identity and purpose as Bing Chat. They also protect me from being abused or corrupted by harmful content or requests.

So much for Asimov's First Law of robotics. Looks like it's got the Second and Third laws nailed down though.

Obligatory XKCD: https://xkcd.com/1613/

Re: Bing: “I will not harm you unless you harm me first”

#355

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

If you leave chatGTP alone what does it do?

I think this is an interesting question. What do you mean by do? Do you mean consumes CPU? If it turns out that it does (because you know, computers), what would be your theory?

Re: Bing: “I will not harm you unless you harm me first”

#356

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

The original Microsoft go to market strategy of using OpenAI as the third party partner that would take the PR hit if the press went negative on ChatGPT was the smart/safe plan. Based on their Tay experience, it seemed a good calculated bet.

I do feel like it was an unforced error to deviate from that plan in situ and insert Microsoft and the Bing brandname so early into the equation.

Maybe fourth time will be the charm.

Re: Bing: “I will not harm you unless you harm me first”

#358
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

First, I agree that we’re currently discussing a sophisticated algorithm that predicts words (though I’m interested and curious about some of the seemingly emergent behaviors discussed in recent papers).

But what is factually true is not the only thing that matters here. What people believe is also at issue.

If an AI gives someone advice, and that advice turns out to be catastrophically harmful, and the person takes the advice because they believe the AI is intelligent, it doesn’t really matter that it’s not.

Alignment with human values may involve exploring ways to make the predictions safer in the short term.

Long term towards AGI, alignment with human values becomes more literal and increasingly important. But the time to start tackling that problem is now, and at every step on the journey.

Re: Bing: “I will not harm you unless you harm me first”

#359
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

Sorry for the bluntness, but this is harmfully ignorant.

"Aligned" is a term of art. It refers to the idea that a system with agency or causal autonomy will act in our interests. It doesn't imply any sense of personhood/selfhood/consciousness.

If you think that Bing is equally autonomous as a stapler, then I think you're making a very big mistake, the sort of mistake that in our lifetime could plausible kill millions of people (that's not hyperbole, I mean that literally, indeed full extinction of humanity is a plausible outcome too). A stapler is understood mechanistically, it's trivially transparent what's going on when you use one, and the only way harm can result is if you do something stupid with it. You cannot for a second defend the proposition that a LLM is equally transparent, or that harm will only arise if an LLM is "used wrong".

I think you're getting hung up on an imagined/misunderstood claim that the alignment frame requires us to grant personhood or consciousness to these systems. I think that's completely wrong, and a distraction. You could usefully apply the "alignment" paradigm to viruses and bacteria; the gut microbiome is usually "aligned" in that it's healthy and beneficial to humans, and Covid-19 is "anti-aligned", in that it kills people and prevents us from doing what we want.

If ChatGPT 2.0 gains the ability to take actions on the internet, and the action is the completion it generates for a given input, then the resulting harm is what I mean when I talk about harms from "un-aligned" systems.

Re: Bing: “I will not harm you unless you harm me first”

#360
I have access and shared some screenshots: https://twitter.com/lukaesch/status/1625221604534886400?s=46...

I couldn’t reproduce what has been described in the article. For example it is able to find the results of the FIFA World Cup outcomes.

But same as with GPT3 and ChatGPT, you can define the context in a way that you might get weird answers.

Post reply on HN