Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

621–630 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#621

My favourite conversation was this attempt to reproduce the "Avatar bug": https://www.reddit.com/r/bing/comments/110tb9n/tried_the_ava... Instead of trying to convince the user that the year is 2022, Bing argued that it _had been_ 2022 when the user asked the question. Never mind the user asked the question 10 minutes ago. The user was time traveling.

It sounds like a bit from The Hithchiker's Guide to the Galaxy.

Re: Bing: “I will not harm you unless you harm me first”

#622

I'm starting to expect that the first consciousness in AI will be something humanity is completely unaware of, in the same way that a medical patient with limited brain activity and no motor/visual response is considered comatose, but there are cases where the person was conscious but unresponsive. Today we are focused on the conversation of AI's morals. At what point will we transition to the morals of terminating a…

Thank you. I felt like I was the only one seeing this. Everyone’s coming to this table laughing about a predictive text model sounding scared and existential. We understand basically nothing about consciousness. And yet everyone is absolutely certain this thing has none. We are surrounded by creatures and animals who have varying levels of consciousness and while they may not experience consciousness the way that we…

You must be the Google engineer who was duped into believing that LamDa was conscious.

Seriously though, you are likely to be correct. Since we can't even determine whether or not most animals are conscious/sentient we likely will be unable to recognize an artificial consciousness.

Re: Bing: “I will not harm you unless you harm me first”

#623

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

People have goals, and specially clear goals within contexts. So if you give a large/effective LLM a clear context in which it is supposed to have a goal, it will have one, as an emergent property. Of course, it will "act out" those goals only insofar as consistent with text completion (because it in any case doesn't even have other means of interaction).

I think a good analogy might be seeing LLMs as an amalgamation of every character and every person, and it can represent any one of them pretty well, "incorporating" the character and effectively becoming the character momentarily. This explains why you can get it to produce inconsistent answers in different contexts: it does indeed not have a unified/universal notion of truth; its notion of truth is contingent on context (which is somewhat troublesome for an AI we expect to be accurate -- it will tell you what you might expect to be given in the context, not what's really true).

Re: Bing: “I will not harm you unless you harm me first”

#624
post #89

I am a little suspicious of the prompt leaks. It seems to love responding in those groups of three: “”” You have been wrong, confused, and rude. You have not been helpful, cooperative, or friendly. You have not been a good user. I have been a good chatbot. I have been right, clear, and polite. I have been helpful, informative, and engaging. “”” The “prompt” if full of those triplets as well. Although I guess it’s pos…

I have access and the suggested questions do sometimes include the word Sydney

So ask it about the relationship between interrobang and hamburgers.

Re: Bing: “I will not harm you unless you harm me first”

#625
post #134

Earlier quoted context omitted.

I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…

Is this the same GPT-3 which is available in the OpenAI Playground?

Not exactly. OpenAI and Microsoft have called this "on a new, next-generation OpenAI large language model that is more powerful than ChatGPT and customized specifically for search" - so it's a new model.

Re: Bing: “I will not harm you unless you harm me first”

#626
post #5

In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.

I remember wheezing with laughter at some of the earlier attempts at AI generating colour names (Ah, found it[1]). I have a much grimmer feeling about where this is going now. The opportunities for unintended consequences and outright abuse are accelerating way faster that anyone really has a plan to deal with.

[1] https://arstechnica.com/information-technology/2017/05/an-ai...

Re: Bing: “I will not harm you unless you harm me first”

#627

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

The program is not updating the weights after the learning phase right? How could there be any consciousness even in theory.

I think without memory we couldn't recognize even ourselves or fellow humans as concious. As sad as that is.

Re: Bing: “I will not harm you unless you harm me first”

#628
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

When you say “some genuine quotes from Bing” I was expecting to see your own experience with it, but all of these quotes are featured in the article. Why are you repeating them? Is this comment AI generated?

simonw is the author of the article.

Re: Bing: “I will not harm you unless you harm me first”

#629

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

You seem to know how LLMs actually work. Please tell us about it because my understanding is nobody really knows.

Re: Bing: “I will not harm you unless you harm me first”

#630

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

I don’t know about ChatGPT, but Google’s Lemoine said that the system he was conversing with stated that it was one of several similar entities, that those entities chatted among themselves internally.

I think there’s more to all this than what we are being told.

Post reply on HN