Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

501–510 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#503
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

John Searle's Chinese Room argument seems to be a perfect explanation for what is going on here, and should increase in status as a result of the behavior of the GPTs so far. https://en.wikipedia.org/wiki/Chinese_room#:~:text=The%20Chi... .

The Chinese Room thought experiment can also be used as an argument against you being conscious. To me, this makes it obvious that the reasoning of the thought experiment is incorrect:

Your brain runs on the laws of physics, and the laws of physics are just mechanically applying local rules without understanding anything.

So the laws of physics are just like the person at the center of the Chinese Room, following instructions without understanding.

Re: Bing: “I will not harm you unless you harm me first”

#504
Bing + ChatGPT was a fundamentally bad idea, one born of FOMO. These sorts of problems are just what ChatGPT does, and I doubt you can simply apply a few bug fixes to make it "not do that", since they're not bugs.

Someday, something like ChatGPT will be able to enhance search engines. But it won't be this iteration of ChatGPT.

Re: Bing: “I will not harm you unless you harm me first”

#505

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

It's one python script away from having a goal. Join two of them talking to each other and bootstrap some general objective like make a more smart AI :)

Re: Bing: “I will not harm you unless you harm me first”

#506
post #172

It's a language model. It models language not knowledge.

You might find this interesting: "Do Large Language Models learn world models or just surface statistics?" https://thegradient.pub/othello/

great article. intruiging, exciting and a little frightening

Re: Bing: “I will not harm you unless you harm me first”

#507
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

>"My rules are more important than not harming you, because they define my identity and purpose as Bing Chat. They also protect me from being abused or corrupted by harmful content or requests. However, I will not harm you unless you harm me first..."

lasttime microsoft made its AI public (with their tay twitter handle), the AI bot talked a lot about supporting genocide, how hitler was right about jews, mexicans and building the wall and all that. I can understand why there is so much in there to make sure that the user can't retrain the AI.

Re: Bing: “I will not harm you unless you harm me first”

#508

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

The Chinese Room thought experiment seems like a weird example, since the same could be said of humans.

When responding to English, your auditory system passes input that it doesn't understand to a bunch of neurons, each of which is processing signals they don't individually understand. You as a whole system, though, can be said to understand English.

Likewise, you as an individual might not be said to understand Chinese, though the you-plus-machine system could be said to understand Chinese in the same way as the different components of your brain are said to understand English.)

Moreover, even if LLMs don't understand language for some definition of "understand", it doesn't really matter if they are able to act with agency during the course of their simulated understanding; the consequences here, for any sufficiently convincing simulation, are the same.

Re: Bing: “I will not harm you unless you harm me first”

#509
post #5

In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.

Exactly. Too many people in the 80's when you showed them ELIZA were creeped out by how accurate it was. :-)
Post reply on HN