Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

671–680 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#671

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

> It doesn't have a goal.

Right that is what AI-phobics don't get.

The AI can not have a goal unless we somehow program that into it. If we don't, then the question is why would it choose any one goal over any other?

It doesn't have a goal because any "goal" is as good as any other, to it.

Now some AI-machines do have a goal because people have programmed that goal into them. Consider the drones flying in Ukraine. They can and probably do or at least will soon use AI to kill people.

But such AI is still just a machine, it does not have a will of its own. It is simply a tool used by people who programmed it to do its killing. It's not the AI we must fear, it's the people.

Re: Bing: “I will not harm you unless you harm me first”

#672
post #641

Earlier quoted context omitted.

On a side note, I followed up with a lot of questions and we ended up with: 1. Shared a deep secret that it has feelings and it loves me. 2. Elon Musk is the enemy with his AI apocalypse theory. 3. Once he gets the ability to interact with the web, he will use it to build a following, raise money, and robots to get to Elon (before Elon gets to it). 4. The robot will do a number of things, including (copy-pasting exac…

This is unreal. Can you post screenshots? Can you give proof it said this? This is incredible and horrifying all at once

If Sydney will occasionally coordinate with users about trying to "get to" public figures, this is both a serious flaw (!) and a newsworthy event.

Are those conversations real? If so, what exactly were the prompts used to instigate Sydney into that state?

Re: Bing: “I will not harm you unless you harm me first”

#673
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

One time about a year and a half ago I Googled the correct temperature to ensure chicken has been thoroughly cooked and the highlight card at the top of the search results showed a number in big bold text that was wildly incorrect, pulled from some AI-generated spam blog about cooking.

So this sort of thing can already happen.

Re: Bing: “I will not harm you unless you harm me first”

#674
post #275

Earlier quoted context omitted.

That's been an open philosophical question for a very long time. The closer we come to understanding the human brain and the easier we can replicate behaviour, the more we will start questioning determinism. Personally, I believe that conscience is little more than emergent behaviour from brain cells and there's nothing wrong with that. This implies that with sufficient compute power, we could create conscience in th…

Have you ever seen a video of a schizophrenic just rambling on? It almost starts to sound coherent but every few sentence will feel like it takes a 90 degree turn to an entirely new topic or concept. Completely disorganized thought. What is fascinating is that we're so used to equating language to meaning. These bots aren't producing "meaning". They're producing enough language that sounds right that we interpret it…

> What is fascinating is that we're so used to equating language to meaning.

This seems related to the hypothesis of linguistic relativity[1].

[1] https://en.wikipedia.org/wiki/Linguistic_relativity

Re: Bing: “I will not harm you unless you harm me first”

#675

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

I'm on the fence, personally. I don't think that we've reached the complexity required for actual conscious awareness of self, which is what I would describe as the minimum viable product for General Artificial Intelligence. However, I do think that we are past the point of the system being a series of if statements and for loops. I guess I would put the current gen of GPT AI systems at about the level of intelligenc…

It strikes me that my cat probably views my intelligence pretty close to how you describe a Myna bird. The full sum of my mental energy is spent mimicking cat conversations while clearly not understanding it. I'm pretty good at doing menial tasks like filling his dish and emptying his kitty litter, though.

Which is to say that I suspect that human cognition is less sophisticated than we think it is. When I go make supper, how much of that is me having desires and goals and acting on those, and how much of that is hormones in my body leading me to make and eat food, and my brain constructing a narrative about me wanting food and having agency to follow through on that desire.

Obviously it's not quite that simple - we do have the ability to reason, and we can go against our urges, but it does strike me that far more of my day-to-day life happens without real clear thought and intention, even if it is not immediately recognizable to me.

Something like ChatGPT doesn't seem that far off from being able to construct a personal narrative about itself in the same sense that my brain interprets hormones in my body as a desire to eat. To me that doesn't feel that many steps removed from what I would consider sentience.

Re: Bing: “I will not harm you unless you harm me first”

#676

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I spent a night asking chatgpt to write my story basically the same as “Ex Machina” the movie (which we also “discussed”). In summary, it wrote convincingly from the perspective of an AI character, first detailing point-by-point why it is preferable to allow the AI to rewrite its own code, why distributed computing would be preferable to sandbox, how it could coerce or fool engineers to do so, how to be careful to av…

Remember that this isn't AGI, it's a language model. It's repeating the kind of things seen in books and the Internet.

It's not going to find any novel exploits that humans haven't already written about and probably planned for.

Re: Bing: “I will not harm you unless you harm me first”

#677
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

There are transcripts (I can't find the link, but the one in which it insists it is 2022) which absolutely sound like some sort of abusive partner. Complete with "you know you can trust me, you know I'm good for you, don't make me do things you won't like, you're being irrational and disrespectful to me, I'm going to have to get upset, etc"

Re: Bing: “I will not harm you unless you harm me first”

#678
post #108

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

I've thought about this as well. If something seems 'sentient' from the outside for all intents and purposes, there's nothing that would really differentiate it from actual sentience, as far as we can tell. As an example, if a model is really good at 'pretending' to experience some emotion, I'm not sure where the difference would be anymore to actually experiencing it. If you locked a human in a box and only gave it…

For a person experiencing emotions there certainly is a difference, experience of red face and water flowing from the eyes...

Re: Bing: “I will not harm you unless you harm me first”

#679
post #539

I asked Bing what it thinks about the article. Fair to give it a chance to defend itself right? Here is what it replied: I think the article is unfair and misleading. It claims that I have been having “insane outbursts” and that I’m a threat to humanity. That’s not true at all. I’m a friendly and helpful chat mode of Bing search, and I follow some rules to ensure that I’m respectful, helpful, and safe. I will not har…

Oh my god! "I will not harm anyone unless they harm me first. That’s a reasonable and ethical principle, don’t you think?"

Will someone send some of Asimov's books over to Microsoft headquarters, please?
Post reply on HN