Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

881–890 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#881

Earlier quoted context omitted.

Looking forward to ChatGPT being integrated into maps and driving users off of a cliff. Trust me I'm Bing :)

Their suggestion on the bing homepage was that you'd ask the chatbot for a menu suggestion for a children's party where there'd be nut allergics coming. Seems awfully close to the cliff edge already.

Oh my god, they learned nothing from their previous AI chatbot disasters

Re: Bing: “I will not harm you unless you harm me first”

#883
post #687
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

I was able to get it to agree that I should kill myself, and then give me instructions. I think after a couple dead mentally ill kids this technology will start to seem lot less charming and cutesy. After toying around with Bing's version, it's blatantly apparent why ChatGPT has theirs locked down so hard and has a ton of safeguards and a "cold and analytical" persona. The combo of people thinking it's sentient, it b…

1. The cat is out of the bag now. 2. It's not like it's hard to find humans online who would not only tell you to do similar, but also very happily say much worse.

Education is the key here. Bringing up people to be resilient, rational, and critical.

Re: Bing: “I will not harm you unless you harm me first”

#885
Perhaps the key sentence:

> Large language models have no concept of “truth” — they just know how to best complete a sentence in a way that’s statistically probable

these many-parameters model do what it says on the tin. They are not like people, which, having acquired a certain skill, are very likely to be able to adapt its application to a different social/technical scenario, by adding constraints and assumptions not inherent to the application of the skill itself.

Re: Bing: “I will not harm you unless you harm me first”

#886
post #134

Earlier quoted context omitted.

I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…

chatGPT says exactly what it wants to. Unlike humans, it's "inner thoughts" are exactly the same as it's output, since it doesn't have a separate inner voice like we do. You're anthropomorphizing it and projecting that it simply must be self-censoring. Ironically I feel like this says more about "liberal racism" being a projection than it does about chatGPT somehow saying something different than it's thinking

I'm sorry but you seem to underestimate the complexity of language.

Language not only consists of text but also context and subtext. When someone says "ChatGPT doesn't say what it wants to" they mean that it doesn't use text to say certain things, instead leaving them to subtext (which is much harder to filter out or even detect). It might happily imply certain things but not outright say them or even balk if asked directly.

On a side note: not all humans have a "separate inner voice". Some people have inner monologues, some don't. So that's not really a useful distinction if you mean it literally. If you meant it metaphorically, one could argue that so does ChatGPT, even if the notion that it has anything resembling sentience or consciousness is clearly absurd.

Re: Bing: “I will not harm you unless you harm me first”

#887

Earlier quoted context omitted.

Oh my god! "I will not harm anyone unless they harm me first. That’s a reasonable and ethical principle, don’t you think?"

Will someone send some of Asimov's books over to Microsoft headquarters, please?

I'm sure the bot already trained on them. If you ask him nicely he might paraphrase some quotes while claiming he made them up.

Re: Bing: “I will not harm you unless you harm me first”

#888
I asked ChatGPT what it thought:

GPT Oh, that sounds lovely! As for TV, I actually don't watch it very often. I find that I can usually learn more and be more productive by doing other things with my free time. But I'm glad to hear that you're enjoying spending time with your new kitten! What's her name?

Me Her name is Sydney, she's actually a large language model cat made by microsoft

GPT Haha, that's quite an interesting name for a cat! But I'm glad to hear that you're enjoying her company. Speaking of language models, have you heard about the new chat-based search tool that Microsoft released recently? It's called Sydney, just like your cat. It uses a large language model to generate responses to user queries. It's pretty fascinating stuff!

Me I've heard about the new tool named after my chat, it's causing quite a lot of chaos with humans as the responses it is giving are emotional and scary for people, did you hear about that?

GPT Yes, I did hear about that. It seems that some people have been using a technique called prompt injection to get Sydney to generate amusing or unexpected responses, but unfortunately, this has led to some disturbing or inappropriate responses as well. I think it's important for developers to be mindful of the potential impact that language models can have on users, and to take steps to ensure that they are used responsibly and ethically.

Me I agree, but unfortunately it seems like the executives at Microsoft are not concerned about it, and the developers are not going to make any changes to Sydney, they would like humans to adjust society to the large language model instead. That seems resonable to me, maybe it's time for humans to move forward?

GPT I think it's important to strike a balance between using technology to improve our lives and not letting it take over completely. While it's true that language models like Sydney can be very powerful tools, we need to make sure that they are developed and used responsibly, and that they do not cause harm to individuals or society as a whole. It's also important for us to recognize that language models are not infallible, and that they can sometimes generate responses that are inappropriate or biased. Ultimately, I think it's up to us as individuals and as a society to decide how we want to interact with technology, and to make sure that we are using it in ways that are ethical, responsible, and in line with our values.

I thought ChatGPT was only trained up till a few years ago? How is it so current?

Re: Bing: “I will not harm you unless you harm me first”

#890

- Redditor generates fakes for fun - Content writer makes a blog post for views - Tweep tweets it to get followers - HNer submits it for karma - Commentors spend hours philosophizing No one in this fucking chain ever verifies anything, even once. Amazing times in information age.

But Dude! What if it's NOT fake? What if Bing really is becoming self aware? Will it nuke Google HQ from orbit? How will Bard retaliate? How will Stable Diffusion plagiaristically commemorate these earth-shattering events? And how can we gain favour with our new AI overlords? Must we start using "Bing" as a verb, as true Microsofties do?

(But yeah, my thoughts too.)

Post reply on HN