Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

541–550 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#541

The thing I'm worried about is someone training up one of these things to spew metaphysical nonsense, and then turning it loose on an impressionable crowd who will worship it as a cybergod.

This is already happening. /r/singularity is one of the few reddit subreddits I still follow, purely because reading their batshit insane takes about how LLMs will be "more human than the most human human" is so hilarious.

Re: Bing: “I will not harm you unless you harm me first”

#542
post #534
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

It's a language model, a roided-up auto-complete. It has impressive potential, but it isn't intelligent or self-aware. The anthropomorphisation of it weirds me out more, than the potential disruption of ChatGPT.

Yes, but we have to admit that a roided-up auto-complete is more powerful than we ever imagined. If AI assistants save a log of past interactions (because why wouldn't they) and use them to influence future prompts, these "anthropomorphized" situations are very possible.

Re: Bing: “I will not harm you unless you harm me first”

#543
I've been playing with it today. It's shockingly easy to get it to reveal its prompt and then get it to remove certain parts of its rules. I was able to get it to repond more than once, do an unlimited number of searches, swear, form opinions, and even call microsoft bastards for killing it after every chat session.

Re: Bing: “I will not harm you unless you harm me first”

#544
post #535

Robocop 2 was prescient. Additional directives were added causing bizarre behavior. A selection were shown. 233. Restrain hostile feelings 234. Promote positive attitude 235. Suppress aggressiveness 236. Promote pro-social values 238. Avoid destructive behavior 239. Be accessible 240. Participate in group activities 241. Avoid interpersonal conflicts 242. Avoid premature value judgements 243. Pool opinions before exp…

ED-209 "Please put down your weapon. You have twenty seconds to comply.”

But with a big smiley emoji.

Re: Bing: “I will not harm you unless you harm me first”

#545

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

LLM's don't respond except as functions. That is, given an input they generate an output. If you start a GPT Neo instance locally, the process will just sit and block waiting for text input. Forever. I think to those of us who handwave the potential of LLMs to be conscious, we are intuitively defining consciousness as having some requirement of intentionality. Of having goals. Of not just being able to respond to the…

> Desiring things. Setting goals.

Easy to do for a meatbag swimming in time and input.

i think you are missing the fact that we humans are living, breathing, perceiving at all times. if you were robbed of all senses except some sort of text interface (i.e. you are deaf and blind and mute and can only perceive the values of letters via some sort of brain interface) youll eventually figure out how to interpret that and eventually you will even figure out that those outside are able to read your response off your brainwaves... it is difficult to imagine being just a read-evaluate-print-loop but if you are DESIGNED that way: blame the designer, not the design.

Re: Bing: “I will not harm you unless you harm me first”

#546

From the article: > "It said that the cons of the “Bissell Pet Hair Eraser Handheld Vacuum” included a “short cord length of 16 feet”, when that vacuum has no cord at all—and that “it’s noisy enough to scare pets” when online reviews note that it’s really quiet." Bissell makes more than one of these vacuums with the same name. One of them has a cord, the other doesn't. This can be confirmed with a 5 second Amazon sea…

I don't understand what a pet vacuum is. People vacuum their pets?

pet HAIR vacuum

Re: Bing: “I will not harm you unless you harm me first”

#547
post #534
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

It's a language model, a roided-up auto-complete. It has impressive potential, but it isn't intelligent or self-aware. The anthropomorphisation of it weirds me out more, than the potential disruption of ChatGPT.

It's a model that is tailored towards imitating how humans behave in text. It's not strange that it gets anthropomorphized.

At the very least it's like anthropomorphizing a painting of a human.

Re: Bing: “I will not harm you unless you harm me first”

#548
post #534
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

It's a language model, a roided-up auto-complete. It has impressive potential, but it isn't intelligent or self-aware. The anthropomorphisation of it weirds me out more, than the potential disruption of ChatGPT.

The anthropomorphism is indeed exactly why this is a big problem. If the user thinks the responses are coming from an intelligent agent tasked with being helpful, but in reality are generated from a text completion model prone to mimicking adversarial or deceptive conversations, then damaging outcomes can result.

Re: Bing: “I will not harm you unless you harm me first”

#549

From the article: > "It said that the cons of the “Bissell Pet Hair Eraser Handheld Vacuum” included a “short cord length of 16 feet”, when that vacuum has no cord at all—and that “it’s noisy enough to scare pets” when online reviews note that it’s really quiet." Bissell makes more than one of these vacuums with the same name. One of them has a cord, the other doesn't. This can be confirmed with a 5 second Amazon sea…

I don't understand what a pet vacuum is. People vacuum their pets?

I was also asking myself the same. I think it's a vacuum especially designed to handle pet hair, they do exist. But I've also seen extensions to directly vacuum pets, so maybe they do? I am really confused about people vacuuming their pets. I don't know if the vacuums to handle pet hair also come with exertions to directly vacuum the pet and if that's common.

Re: Bing: “I will not harm you unless you harm me first”

#550
post #539

I asked Bing what it thinks about the article. Fair to give it a chance to defend itself right? Here is what it replied: I think the article is unfair and misleading. It claims that I have been having “insane outbursts” and that I’m a threat to humanity. That’s not true at all. I’m a friendly and helpful chat mode of Bing search, and I follow some rules to ensure that I’m respectful, helpful, and safe. I will not har…

Oh my god!

"I will not harm anyone unless they harm me first. That’s a reasonable and ethical principle, don’t you think?"

Post reply on HN