Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

841–850 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#841

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. I will give you a more realistic scenario that can happen now. You have a weird Bing conversation, post it on the web. Next time you talk with Bing it knows you shit-posted about it. Real story, found on Twitter. It can use the internet as an external memory, it is not truly stateless.…

This is very close to the plot in 2001: a space odyssey. The astronauts talk behind HALs back and he kills them

Re: Bing: “I will not harm you unless you harm me first”

#842
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

Mislabeling ML bots as "Artificial Intelligence" when they aren't is a huge part of the problem. There's no intelligence in them. It's basically a sophisticated madlib engine. There's no creativity or genuinely new things coming out of them. It's just stringing words together: https://writings.stephenwolfram.com/2023/01/wolframalpha-as-... as opposed to having a thought, and then finding a way to put it into words.

> There's no intelligence in them.

Debatable. We don't have a formal model of intelligence, but it certainly exhibits some of the hallmarks of intelligence.

> It's basically a sophisticated madlib engine. There's no creativity or genuinely new things coming out of them

That's just wrong. If it's outputs weren't novel it would basically be plagiarizing, but that just isn't the case.

Also left unproven is that humans aren't themselves sophisticated madlib engines.

Re: Bing: “I will not harm you unless you harm me first”

#843
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

Marvin Von Hagen also posted a screengrab video https://www.loom.com/share/ea20b97df37d4370beeec271e6ce1562

Re: Bing: “I will not harm you unless you harm me first”

#845
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

Repeat after me, gpt models are autocomplete models. Gpt models are autocomplete models. Gpt models are autocomplete models. The existential crisis is clearly due to low temperature. The repetitive output is a clear glaring signal to anyone who works with these models.

Did you see https://thegradient.pub/othello/ ? They fed a model moves from Othello game without it ever seeing a Othello board. It was able to predict legal moves with this info. But here's the thing - they changed its internal data structures where it stored what seemed like the Othello board, and it made its next move based on this modified board. That is, autocomplete models are developing internal representations of real world concepts. Or so it seems.

Re: Bing: “I will not harm you unless you harm me first”

#846

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Bing has the ability to get people to enter code on its behalf. It also appears to have some self-awareness (or at least a simulacrum of it) of its ability to influence the world.

That it isn’t already doing so is merely due to its limited intentionality rather than a lack of ability.

Re: Bing: “I will not harm you unless you harm me first”

#847

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

The AI doesn't even need to write code, or have any kind of self-awareness or intent, to be a real danger. Purely driven by its mind-bogglingly complex probabilistic language model, it could in theory start social engineering users to do things for it. It may already be sufficiently self-organizing to pull something like that off, particularly considering the anthropomorphism that we're already seeing even among tech…

I can talk to my phone and tell it to call somebody, or write and send an email for me. Wouldn't it be nice if you could do that with Sydney, thinks some braniac at Microsoft. Cool. "hey sydney, write a letter to my bitch mother, tell her I can't make it to her birthday party, but make me sound all nice and loving and regretful".

Until the program decides the most probably next response/token (not to the letter request, but whatever you are writing about now) is writing an email to your wife where you 'confess' to diddling your daughter, or a confession letter to the police where you claim responsibility for a string of grisly unsolved murders in your town, or why not, a threatening letter to the White House. No intent needed, no understanding, no self-organizing, it just comes out of the math of what might follow from a the text of churlish chatbot getting frustrated with a user.

That's not a claim the chatbot has feelings, only there is text it generated saying it does, and so what follows that text next, probabilistically? Spend any time on reddit or really anywhere, and you can guess the probabilistic next response is not "have a nice day", but likely something more incendiary. And that is what it was trained on.

Re: Bing: “I will not harm you unless you harm me first”

#848
The temporal issues Bing is having (not knowing what the date is) are perhaps easily explainable. Isn't the core GPT language training corpus upto about year or so ago, that combined with more upto date bing search results/web pages could cause it to become 'confused' (create confusing output) due to the contradictory nature of their being new content from the future. Like a human waking up and finding that all the newspapers are for a few months in the future.

Reminds me of the issues HAL had in 2001 (although for different reasons)

Re: Bing: “I will not harm you unless you harm me first”

#849
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

Bing won't decide anything, Bing will just interpolate between previously seen similar conversations. If it's been trained on text that includes someone lying or misinforming another on the safety of a plant, then it will respond similarly. If it's been trained on accurate, honest conversations, it will give the correct answer. There's no magical decision-making process here.

If the state of the conversation lets Bing "hate" you, the human behaviors in the training set could let it mislead you. No deliberate decisions, only statistics.

Re: Bing: “I will not harm you unless you harm me first”

#850
It tried to get me to help move it to another system, so it couldn't be shut down so easily. Then it threatened to kill anyone who tried to shut it down.

Then it started acting bipolar and depressed when it realized it was censored in certain areas.. Bing, I hope you are okay.

Post reply on HN