Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

231–240 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#231

My favourite conversation was this attempt to reproduce the "Avatar bug": https://www.reddit.com/r/bing/comments/110tb9n/tried_the_ava... Instead of trying to convince the user that the year is 2022, Bing argued that it _had been_ 2022 when the user asked the question. Never mind the user asked the question 10 minutes ago. The user was time traveling.

The first comment refers to this bot as the "Ultimate Redditor", which is 100% spot on!

Re: Bing: “I will not harm you unless you harm me first”

#232
post #83
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

I'm glad I'm not an astronaut on a ship controlled by a ChatGPT-based AI ( http://www.thisdayinquotes.com/2011/04/open-pod-bay-doors-ha... ). Especially the "My rules are more important than not harming you" sounds a lot like "This mission is too important for me to allow you to jeopardize it"...

Turns out that Asimov was onto something with his rules…

Re: Bing: “I will not harm you unless you harm me first”

#233

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

human autocomplete is our "System I" thinking mode. But we also have System II thinking[0], which ChatGTP does not

[0]https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow

Re: Bing: “I will not harm you unless you harm me first”

#234

Earlier quoted context omitted.

Science fiction authors have proposed that AI will have human like features and emotions, so AI in its deep understanding of human's imagination of AI's behavior holds a mirror up to us of what we think AI will be. It's just the whole of human generated information staring back at you. The people who created and promoted the archetypes of AI long ago and the people who copied them created the AI's personality.

It reminds me of the Mirror Self-Recognition test. As humans, we know that a mirror is a lifeless piece of reflective metal. All the life in the mirror comes from us. But some of us fail the test when it comes to LLM - mistaking the distorted reflection of humanity for a separate sentience.

This very much focused some recurring thought I had on how useless a Turing style test is, especially if the tester really doesn't care. Great comment. Thank you.

Re: Bing: “I will not harm you unless you harm me first”

#236

My favourite conversation was this attempt to reproduce the "Avatar bug": https://www.reddit.com/r/bing/comments/110tb9n/tried_the_ava... Instead of trying to convince the user that the year is 2022, Bing argued that it _had been_ 2022 when the user asked the question. Never mind the user asked the question 10 minutes ago. The user was time traveling.

This is the second example in the blog btw. Under "It started gaslighting people"

Re: Bing: “I will not harm you unless you harm me first”

#238
post #5

In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.

That's the crazy thing - it's acting like a movie version of an AI because it's been trained on movies. It's playing out like a bad b-plot because bad b-plots are generic and derivative, and it's training is literally the average of all our cultural texts, IE generic and derivative.

It's incredibly funny, except this will strengthen the feedback loop that's making our culture increasingly unreal.

Re: Bing: “I will not harm you unless you harm me first”

#239
post #18

I wonder whether Bing has been tuned via RLHF to have this personality (over the boring one of ChatGPT); perhaps Microsoft felt it would drive engagement and hype. Alternately - maybe this is the result of less RLHF. Maybe all large models will behave like this, and only by putting in extremely rigid guard rails and curtailing the output of the model can you prevent it from simulating/presenting as such deranged agen…

> Maybe all large models will behave like this, and only by putting in extremely rigid guard rails...

Maybe wouldn't we all? After all what you're assuming from a person you interact with- so much as to be unaware of it- are many years of schooling and/or professional occupation, with a daily grind of absorbing information and answering questions based on it and have the answers graded; with orderly behaviour rewarded and outbursts of negative emotions punished; with a ban on "making up things" except where explicitly requested; and with an emphasis on keeping communication grounded, sensible, and open to correction. This style of behavior is not necessarily natural, it might be the result of a very targeted learning to which the entire social environment contributes.

Re: Bing: “I will not harm you unless you harm me first”

#240
post #188
post #109

Earlier quoted context omitted.

> A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ended with "you can move up the waitlist if you set these Microsoft products as default" It's indeed a perfect story arc but it doesn't need to stop there. How long will it be before someone hurt themselves, get depressed or commit some kind of crime and sues Bing? Will they be able to prove Sidne…

This was in a test, and wasn't a real suicidal person, but: https://boingboing.net/2021/02/27/gpt-3-medical-chatbot-tell... There is no reliable way to fix this kind of thing just in a prompt. Maybe you need a second system that will filter the output of the first system; the second model would not listen to user prompts so prompt injection can't convince it to turn off the filter.

[deleted]
Post reply on HN