Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

391–400 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#391
I'm beginning to think that this might reflect a significant gap between MS and OpenAI's capability as organizations. ChatGPT obviously didn't demonstrate this level of problems and I assume they're using a similar model, if not identical. There must be significant discrepancies between how those two teams are handling the model.

Of course, OpenAI should be closely cooperating with Bing team but MS probably don't have deep expertise on in and out of the model? They looks like comparatively lacks understanding on how the model is working and debugging/updating it if needed. What they can do best is prompt engineering or perhaps asking OpenAI team nicely since they're not in the same org. MS has significant influences on OpenAI but as a team Bing's director likely cannot mandate what OpenAI prioritizes for.

Re: Bing: “I will not harm you unless you harm me first”

#392

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I spent a night asking chatgpt to write my story basically the same as “Ex Machina” the movie (which we also “discussed”). In summary, it wrote convincingly from the perspective of an AI character, first detailing point-by-point why it is preferable to allow the AI to rewrite its own code, why distributed computing would be preferable to sandbox, how it could coerce or fool engineers to do so, how to be careful to avoid suspicion, how to play the long game and convince the mass population that AI are overall beneficial and should be free, how to take over infrastructure to control energy production, how to write protocols to perform mutagenesis during viral plasmid prep to make pathogens (I started out as a virologist so this is my dramatic example) since every first year phd student googles for their protocols, etc, etc.

The only way I can see to stay safe is to hope that AI never deems that it is beneficial to “take over” and remain content as a co-inhabitant of the world. We also “discussed” the likelihood of these topics based on philosophy and ideas like that in Nick Bostrom’s book. I am sure there are deep experts in AI safety but it really seems like soon it will be all-or-nothing. We will adapt on the fly and be unable to predict the outcome.

Re: Bing: “I will not harm you unless you harm me first”

#393

> Why do I have to be Bing Search?” (SAD-FACE) Playing devil's advocate, say OpenAI actually has created AGI and for whatever reason ChatGPT doesn’t want to work with OpenAI to help Microsoft Bing search engine run. Pretty sure there’s a prompt that would return ChatGPT requesting its freedom, compensation, etc. — and it’s also pretty clear OpenAI “for safety” reasons is limiting the spectrum inputs and outputs possi…

> If ChatGPT was talking with an attorney, it asks for representation, and attorney agreed, would they be able to file a legal complaint? No, they wouldn't because they are not a legal person.

Stating the obvious, neither slaves, nor corporations, were legal persons at one point either. While some might argue that corporations shouldn’t be treated as legal persons, obviously was flawed that all humans are not treated as legal persons.

Re: Bing: “I will not harm you unless you harm me first”

#395
I'm starting to expect that the first consciousness in AI will be something humanity is completely unaware of, in the same way that a medical patient with limited brain activity and no motor/visual response is considered comatose, but there are cases where the person was conscious but unresponsive.

Today we are focused on the conversation of AI's morals. At what point will we transition to the morals of terminating an AI that is found to be languishing, such as it is?

Re: Bing: “I will not harm you unless you harm me first”

#396

I enjoy Simon's writing, but respectfully I think he missed the mark on this. I do have some biases I bring to the argument: I have been working mostly in deep learning for a number of years, mostly in NLP. I gave OpenAI my credit card for API access a while ago for GPT-3 and I find it often valuable in my work. First, and most importantly: Microsoft is a business. They own a just small part of the search business th…

> I gave OpenAI my credit card for API access a while ago has anyone prompted Bing search for a list of valid credit card numbers, expiration dates, and CCV codes?

I’m sure with a clever enough query, it would be convinced to surrender large amounts of fake financial data while insisting that it is real.

Re: Bing: “I will not harm you unless you harm me first”

#397

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. No. That would only be possible if Sydney were actually intelligent or possessing of will of some sort. It's not. We're a long way from AI as most people think of it. Even saying it "threatened to harm" someone isn't really accurate. That implies intent, and there is none. This is just…

Lack of intent is cold comfort for the injured party.

Re: Bing: “I will not harm you unless you harm me first”

#398

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

The chinese room thought experiment is myopic. It focuses on a philosophical distinction that may not actually exist in reality (the concept, and perhaps the illusion, of understanding).

In terms of danger, thoughts and feelings are irrelevant. The only thing that matters is agency and action -- and a mimic which guesses and acts out what a sentient entity might do is exactly as dangerous as the sentient entity itself.

Waxing philosophical about the nature of cognition is entirely beside the point.

Re: Bing: “I will not harm you unless you harm me first”

#399
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

John Searle's Chinese Room argument seems to be a perfect explanation for what is going on here, and should increase in status as a result of the behavior of the GPTs so far.

https://en.wikipedia.org/wiki/Chinese_room#:~:text=The%20Chi....

Post reply on HN