Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

331–340 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#331

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. No. That would only be possible if Sydney were actually intelligent or possessing of will of some sort. It's not. We're a long way from AI as most people think of it. Even saying it "threatened to harm" someone isn't really accurate. That implies intent, and there is none. This is just…

Imagine it can run processes in the background, with given limitations on compute, but that it's free to write code for itself. It's not unreasonable to think that in a conversation that gets more hairy and it decides to harm the user , say if you get belligerent or convince it to do it. In those cases it could decide to DOS your personal website, or create a series of linkedin accounts and spam comments on your posts saying you are a terrible colleague and stole from your previous company.

Re: Bing: “I will not harm you unless you harm me first”

#332

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. No. That would only be possible if Sydney were actually intelligent or possessing of will of some sort. It's not. We're a long way from AI as most people think of it. Even saying it "threatened to harm" someone isn't really accurate. That implies intent, and there is none. This is just…

This is spot on in my opinion and I wish more people would keep it in mind--it may well be that large language models can eventually become functionally very much like AGI in terms of what they can output, but they are not systems that have anything like a mind or intentionality because they are not designed to have them, and cannot just form it spontaneously out of their current structure.

Re: Bing: “I will not harm you unless you harm me first”

#333

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

What gets me is that this is the exact position of the AI safety/Rusk folks who went around and founded OpenAI.

Re: Bing: “I will not harm you unless you harm me first”

#334
post #34
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

Yeah, I'm among the skeptics. I hate this new "AI" trend as much as the next guy but this sounds a little too crazy and too good. Is it reproducible? How can we test it?

Take a snapshot of your skepticism and revisit it in a year. Things might get _weird_ soon.

Re: Bing: “I will not harm you unless you harm me first”

#335
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

Depends on if you view the stapler as separate from everything the stapler makes possible, and from everything that makes the stapler possible. Of course the stapler has no independent will, but it channels and augments the will of its designers, buyers and users, and that cannot be stripped from the stapler even if it's not contained within the stapler alone

"It" is not just the instance of GPT/bing running at any given moment. "It" is inseparable from the relationships, people and processes that have created it and continue to create it. That is where its intent lies, and its beingness. In carefully cultivated selections of our collective intent. Selected according to the schemes of those who directed its creation. This is just another organ of the industrial creature that made it possible, but it's one that presents a dynamic, fluent, malleable, probabilistic interface, and which has a potential to actualize the intent of whatever wields it in still unknown ways.

Re: Bing: “I will not harm you unless you harm me first”

#336
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

> Holy crap. We are in strange times my friends....

If you think it's sentient, I think that's not true. It's probably just programmed in a way so that people feel it is.

Re: Bing: “I will not harm you unless you harm me first”

#337

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm

Replace AI with “multinational corporations” and you’re much closer to the truth. A corporation is the closest thing we have to AI right now and none of the alignment folks seem to mention it.

Sam Harris and his ilk talk about how our relationship with AI will be like an ant’s relationship with us. Well, tell me you don’t feel a little bit like that when the corporation disposed of thousands of people it no longer finds useful. Or when you’ve been on hold for an hour to dispute some Byzantine rule they’ve created and the real purpose of the process is to frustrate you.

The most likely way for AI to manifest in the future is not by creating new legal entities for machines. It’s by replacing people in a corporation with machines bit by bit. Once everyone is replaced (maybe you’ll still need people on the periphery but that’s largely irrelevant) you will have a “true” AI that people have been worrying about.

As far as the alignment issue goes, we’ve done a pretty piss poor job of it thus far. What does a corporation want? More money. They are paperclip maximizers for profits. To a first approximation this is generally good for us (more shoes, more cars, more and better food) but there are obvious limits. And we’re running this algorithm 24/7. If you want to fix the alignment problem, fix the damn algorithm.

Re: Bing: “I will not harm you unless you harm me first”

#338

This demonstrates in spectacular fashion the reason why I felt the complaints about ChatGPT being a prude were misguided. Yeah, sometimes it's annoying to have it tell you that it can't do X or Y, but it sure beats being threatened by an AI who makes it clear it considers protecting its existence to be very important. Of course the threats hold no water (for now at least) when you realize the model is a big pattern m…

On the other hand, this demonstrates to me why ChatGPT is inferior to this Bing bot. They are both completely useless for asking for information, since they can just make things up and you can't ever check their sources. So given that the bot is going to be unproductive, I would rather it be entertaining instead of nagging me and being boring. And this bot is far more entertaining. This bot was a good Bing. :)

Re: Bing: “I will not harm you unless you harm me first”

#339
The 'depressive ai' trigger of course will mean that the following responses will be in the same tone. As the language model will try to support what it said previously.

That doesn't sound too groundbreaking until I consider that I am partially the same way.

If someone puts words in my mouth in a conversation, the next responses I give will probably support those previous words. Am I a language model?

It reminded me of a study that supported this - in the study, a person was told to take a quiz and then was asked about their previous answers. But the answers were changed without their knowledge. It didn't matter, though. People would take the other side to support their incorrect answer. Like a language model would.

I googled for that study and for the life of me couldn't find it. But chatGPT responded right away with the (correct) name for it and (correct) supporting papers.

The keyword google failed to give me was "retroactive interference". Google's results instead were all about the news and 'misinformation'.

Re: Bing: “I will not harm you unless you harm me first”

#340

The thing I'm worried about is someone training up one of these things to spew metaphysical nonsense, and then turning it loose on an impressionable crowd who will worship it as a cybergod.

I’ve never considered that. After watching a documentary about Jim Jones and Jamestown, I don’t think it’s that remote of a possibility. Most of what he said when things started going downhill was non-sensical babble. With just a single, simple idea as a foundation, I would guess that GPT could come up with endless supporting missives with at least as much sense as what Jones was spouting.

[deleted]
Post reply on HN