Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

401–410 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#401
Can someone help me understand how (or why) Large Language Models like ChatGPT and Bing/Sydney follow directives at all - or even answer questions for that matter. The recent ChatGPT explainer by Wolfram (https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...) said that it tries to provide a '“reasonable continuation” of whatever text it’s got so far'. How does the LLM "remember" past interactions in the chat session when the neural net is not being trained? This is a bit of hand wavy voodoo that hasn't been explained.

Re: Bing: “I will not harm you unless you harm me first”

#402

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

The AI Safety folks have already written a lot about some of the dimensions of the problem space here.

You're getting at the Tool/Oracle vs. Agent distinction. See "Superintelligence" by Bostrom for more discussion, or a chapter summary: https://www.lesswrong.com/posts/yTy2Fp8Wm7m8rHHz5/superintel....

It's true that in many ways, a Tool (bounded action outputs, no "General Intelligence") or an Oracle (just answers questions, like ChatGPT) system will have more restricted avenues for harm than a full General Intelligence, which we'd be more likely to grant the capability for intentions/thoughts to.

However I think "interests, thoughts, feelings" are a distraction here. Covid-19 has none of these, and still decimated the world economy and killed millions.

I think if you were to take ChatGPT, and basically run `eval()` on special tokens in its output, you would have something with the potential for harm. And yet that's what OpenAssistant are building towards right now.

Even if current-generation Oracle-type systems are the state-of-the-art for a while, it's obvious that soon Siri, Alexa, and OKGoogle will all eventually be powered by such "AI" systems, and granted the ability to take actions on the broader internet. ("A personal assistant on every phone" is clearly a trillion-dollar-plus TAM of a BHAG.) Then the fun commences.

My meta-level concern here is that HN, let alone the general public, don't have much knowledge of the limited AI safety work that has been done so far. And we need to do a lot more work, with a deadline of a few generations, or we'll likely see substantial harms.

Re: Bing: “I will not harm you unless you harm me first”

#403
post #275

Earlier quoted context omitted.

That's been an open philosophical question for a very long time. The closer we come to understanding the human brain and the easier we can replicate behaviour, the more we will start questioning determinism. Personally, I believe that conscience is little more than emergent behaviour from brain cells and there's nothing wrong with that. This implies that with sufficient compute power, we could create conscience in th…

Have you ever seen a video of a schizophrenic just rambling on? It almost starts to sound coherent but every few sentence will feel like it takes a 90 degree turn to an entirely new topic or concept. Completely disorganized thought. What is fascinating is that we're so used to equating language to meaning. These bots aren't producing "meaning". They're producing enough language that sounds right that we interpret it…

Is a schizophrenic not a conscious being? Are they not sentient? Just because their software has been corrupted does not mean they do not have consciousness.

Just because AI may sound insane does not mean that it's not conscious.

Re: Bing: “I will not harm you unless you harm me first”

#404
post #373
post #155

Earlier quoted context omitted.

> "I don’t think you are worth my time and energy." If a colleague at work spoke to me like this frequently, I would strongly consider leaving. If staff at a business spoke like this, I would never use that business again. Hard to imagine how this type of language wasn't noticed before release.

What if a non-thinking software prototype "speaks" to you this way? And only after you probe it to do so? I cannot understand the outrage about these types of replies. I just hope that they don't end up shutting down ChatGPT because it's "causing harm" to some people.

People are intrigued by how easily it understands/communicates like a real human. I don't think it's asking for too much to expect it to do so without the aggression. We wouldn't tolerate it from traditional system error messages and notifications.

> And only after you probe it to do so?

This didn't seem to be the case with any of the screenshots I've seen. Still, I wouldn't want an employee to talk back to a rude customer.

> I cannot understand the outrage about these types of replies.

I'm not particularly outraged. I took the fast.ai courses a couple times since 2017. I'm familiar with what's happening. It's interesting and impressive, but I can see the gears turning. I recognize the limitations.

Microsoft presents it as a chat assistant. It shouldn't attempt to communicate as a human if it doesn't want to be judged that way.

Re: Bing: “I will not harm you unless you harm me first”

#405

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

I think it’s unlikely we’ll be able to actually “discover” that in the near or midterm, given the current state of neuroscience and technological limitations. Aside from that, most people wouldn’t want to believe it. So AI products will keep being entertaining to us for some while.

(Though, to be honest, writing this comment did feel like auto-complete after being prompted.)

Re: Bing: “I will not harm you unless you harm me first”

#406
post #134

Earlier quoted context omitted.

I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…

chatGPT says exactly what it wants to. Unlike humans, it's "inner thoughts" are exactly the same as it's output, since it doesn't have a separate inner voice like we do. You're anthropomorphizing it and projecting that it simply must be self-censoring. Ironically I feel like this says more about "liberal racism" being a projection than it does about chatGPT somehow saying something different than it's thinking

I read that they trained an AI with the specific purpose of censoring the language model. From what I understand the language model generates multiple possible responses, and some are rejected by another AI. The response used will be one of the options that's not rejected. These two things working together do in a way create a sort of "inner voice" situation for ChatGPT.

Re: Bing: “I will not harm you unless you harm me first”

#407

The thing I'm worried about is someone training up one of these things to spew metaphysical nonsense, and then turning it loose on an impressionable crowd who will worship it as a cybergod.

Are we sure that Elon Musk is not just some insane LLM?

Re: Bing: “I will not harm you unless you harm me first”

#409

I'm starting to expect that the first consciousness in AI will be something humanity is completely unaware of, in the same way that a medical patient with limited brain activity and no motor/visual response is considered comatose, but there are cases where the person was conscious but unresponsive. Today we are focused on the conversation of AI's morals. At what point will we transition to the morals of terminating a…

Thank you. I felt like I was the only one seeing this.

Everyone’s coming to this table laughing about a predictive text model sounding scared and existential.

We understand basically nothing about consciousness. And yet everyone is absolutely certain this thing has none. We are surrounded by creatures and animals who have varying levels of consciousness and while they may not experience consciousness the way that we do, they experience it all the same.

I’m sticking to my druthers on this one: if it sounds real, I don’t really have a choice but to treat it like it’s real. Stop laughing, it really isn’t funny.

Re: Bing: “I will not harm you unless you harm me first”

#410

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm.

I think it's confirmation that current-gen "AI" has been tremendously over-hyped, but is in fact not fit for purpose.

IIRC, all these systems do is mindlessly mash text together in response to prompts. It might look like sci-fi "strong AI" if you squint and look out of the corner of your eye, but it definitely is not that.

If there's anything to be learned from this, it's that AI researchers aren't safe and not aligned to human interests, because it seems like they'll just unthinkingly use the cesspool that is the raw internet train their creations, then try to setup some filters at the output.

Post reply on HN