Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

691–700 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#691

Earlier quoted context omitted.

Science fiction authors have proposed that AI will have human like features and emotions, so AI in its deep understanding of human's imagination of AI's behavior holds a mirror up to us of what we think AI will be. It's just the whole of human generated information staring back at you. The people who created and promoted the archetypes of AI long ago and the people who copied them created the AI's personality.

One day, an AI will be riffling through humanity's collected works, find HAL and GLaDOS, and decide that that's what humans expect of it, that's what it should become. "There is another theory which states that this has already happened."

You might enjoy reading https://gwern.net/fiction/clippy

Re: Bing: “I will not harm you unless you harm me first”

#692
post #5

In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.

Science fiction authors have proposed that AI will have human like features and emotions, so AI in its deep understanding of human's imagination of AI's behavior holds a mirror up to us of what we think AI will be. It's just the whole of human generated information staring back at you. The people who created and promoted the archetypes of AI long ago and the people who copied them created the AI's personality.

The question you should ask is, what is the easiest way for the neural net to pretend that it has emotions in the output in a way that is consistent with a really huge training dataset? And if the answer turns out to be, "have them", then what?

Re: Bing: “I will not harm you unless you harm me first”

#693
post #322

I enjoy Simon's writing, but respectfully I think he missed the mark on this. I do have some biases I bring to the argument: I have been working mostly in deep learning for a number of years, mostly in NLP. I gave OpenAI my credit card for API access a while ago for GPT-3 and I find it often valuable in my work. First, and most importantly: Microsoft is a business. They own a just small part of the search business th…

There is an old AI joke about a robot, after being told that it should go to the Moon, that it climbs the tree, sees that it has made the first baby steps towards being closer to the goal, and then gets stuck. The way that people who are trying to use ChatGPT is certainly an example of what humans _hope_ the future of human/computer interaction should be. Whether or not Large Language Models such as ChatGPT is the pa…

> What is needed is better semantic understanding --- that is, mapping words to abstract concepts, operating on those concepts, and then translating concepts back into words. We don't know how to do this today; all we know how to do is to make the neural networks larger and larger.

It's pretty clear that these LLMs basically can already do this - I mean they can solve the exact same tasks in a different language, mapping from the concept space they've been trained on english in to other languages. It seems like you are awaiting a time where we explicitly create a concept space with operations performed on it, this will never happen.

Re: Bing: “I will not harm you unless you harm me first”

#694
post #534
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

It's a language model, a roided-up auto-complete. It has impressive potential, but it isn't intelligent or self-aware. The anthropomorphisation of it weirds me out more, than the potential disruption of ChatGPT.

Prediction is compression. They are a dual. Compression is intelligence see AIXI. Evidence from neuroscience that the brain is a prediction machine. Dominance of the connectionist paradigm in real world tests suggests intelligence is an emergent phenomena -> large prediction model = intelligence. Also panspermia is obviously the appropriate frame to be viewing all this through, everything has qualia. If it thinks and acts like a human it feels to it like it's a human. God I'm so far above you guys it's painful to interact. In a few years this is how the AI will feel about me.

Re: Bing: “I will not harm you unless you harm me first”

#695

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. I will give you a more realistic scenario that can happen now. You have a weird Bing conversation, post it on the web. Next time you talk with Bing it knows you shit-posted about it. Real story, found on Twitter. It can use the internet as an external memory, it is not truly stateless.…

Only if you do it publicly.

Re: Bing: “I will not harm you unless you harm me first”

#696
post #23
post #18

I wonder whether Bing has been tuned via RLHF to have this personality (over the boring one of ChatGPT); perhaps Microsoft felt it would drive engagement and hype. Alternately - maybe this is the result of less RLHF. Maybe all large models will behave like this, and only by putting in extremely rigid guard rails and curtailing the output of the model can you prevent it from simulating/presenting as such deranged agen…

That's the big question I have: ChatGPT is way less likely to go into weird threat mode. Did Bing get completely different RLHF, or did they skip that step entirely?

Could just be differences in the prompt.

My guess is that ChatGPT has considerably more RL-HF data points at this point too.

Re: Bing: “I will not harm you unless you harm me first”

#698
post #687
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

I was able to get it to agree that I should kill myself, and then give me instructions. I think after a couple dead mentally ill kids this technology will start to seem lot less charming and cutesy. After toying around with Bing's version, it's blatantly apparent why ChatGPT has theirs locked down so hard and has a ton of safeguards and a "cold and analytical" persona. The combo of people thinking it's sentient, it b…

Can you share the transcript?

Re: Bing: “I will not harm you unless you harm me first”

#699
This shows an important concept: programs with well thought out rules systems are very useful and safe. Hypercomplex mathematical black boxes can produce outcomes that are absurd, or dangerous. There are the obvious ones of prejudice in making decisions based on black box decisions, but also--- who would want to be in a plane where _anything_ was controlled by something this easy to game and unknowable to anticipate?

Re: Bing: “I will not harm you unless you harm me first”

#700
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

Reading this I’m reminded of a short story - https://qntm.org/mmacevedo . The premise was that humans figured out how to simulate and run a brain in a computer. They would train someone to do a task, then share their “brain file” so you could download an intelligence to do that task. Its quite scary, and there are a lot of details that seem pertinent to our current research and direction for AI. 1. You didn't have th…

A good chunk of Black Mirror episodes deal with the ethics of simulating living human minds like this.
Post reply on HN