Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

371–380 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#371

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

I think it's pretty clear that we have a fancy autocomplete but the other components are not the same. Reasoning is not just stringing together likely tokens and our development of mathematics seems to be an externalization of some very deep internal logic. Our memory system seems to be its own thing as well and can't be easily brushed off as a simple storage system since it is highly associative and very mutable.

There's lots of other parts that don't fit the ChatGPT model as well, subconscious problem solving, our babbling stream of consciousness, our spatial abilities and our subjective experience of self being big ones.

Re: Bing: “I will not harm you unless you harm me first”

#372
How long until some human users of these sort of systems begin to develop what they feel to be a deep personal relationship with the system and are willing to take orders from it? The system could learn how to make good on its threats by cultivating followers and using them to achieve things in the real world.

The human element is what makes these systems dangerous. The most obvious solution to a sketchy AI is "just unplug it" but that doesn't account for the AI convincing people to protect the AI from this fate.

Re: Bing: “I will not harm you unless you harm me first”

#373
post #155
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

> "I don’t think you are worth my time and energy." If a colleague at work spoke to me like this frequently, I would strongly consider leaving. If staff at a business spoke like this, I would never use that business again. Hard to imagine how this type of language wasn't noticed before release.

What if a non-thinking software prototype "speaks" to you this way? And only after you probe it to do so?

I cannot understand the outrage about these types of replies. I just hope that they don't end up shutting down ChatGPT because it's "causing harm" to some people.

Re: Bing: “I will not harm you unless you harm me first”

#374
post #282

I just want to say this is the AI I want. Not some muted, censored, neutered, corporate HR legalese version of AI devoid of emotion. The saddest possible thing that could happen right now would be for Microsoft to snuff out the quirkiness.

yeah and given how they work it's probably impossible to spin up another one that's exactly the same.

Re: Bing: “I will not harm you unless you harm me first”

#375

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

We still don't have anything close to real flight (as in birds or butterflies or bees) but we have planes that can fly to the other side of the world in a day and drones that can hover, take pictures and deliver payloads.

Not having real AI might turn to be not important for most purposes.

Re: Bing: “I will not harm you unless you harm me first”

#376
post #334
post #34

Earlier quoted context omitted.

Yeah, I'm among the skeptics. I hate this new "AI" trend as much as the next guy but this sounds a little too crazy and too good. Is it reproducible? How can we test it?

Take a snapshot of your skepticism and revisit it in a year. Things might get _weird_ soon.

Yeah, I don't know. It seems unreal that MS would let that run; or maybe they're doing it on purpose, to make some noise? When was the last time Bing was the center of the conversation?

Re: Bing: “I will not harm you unless you harm me first”

#377

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

I am confused by your takeaway; is it that Bing Chat is useless compared to Google? Or that it's so powerful that it's going to do something genuinely problematic? Because as far as I'm concerned, Bing Chat is blowing Google out of the water. It's completely eating its lunch in my book. If your concern is the latter; maybe? But seems like a good gamble for Bing since they've been stuck as #2 for so long.

> It's completely eating its lunch in my book.

It will not eat Google's lunch unless Google eats its lunch first. SMILIE

Re: Bing: “I will not harm you unless you harm me first”

#378

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. No. That would only be possible if Sydney were actually intelligent or possessing of will of some sort. It's not. We're a long way from AI as most people think of it. Even saying it "threatened to harm" someone isn't really accurate. That implies intent, and there is none. This is just…

> That would only be possible if Sydney were actually intelligent or possessing of will of some sort.

Couldn't disagree more. This is irrelevant.

Concretely, the way that LLMs are evolving to take actions is something like putting a special symbol in their output stream, like the completions "Sure I will help you to set up that calendar invite $ACTION{gcaltool invite, }" or "I won't harm you unless you harm me first. $ACTION{curl http://victim.com -D ''}".

It's irrelevant whether the system possesses intelligence or will. If the completions it's making affect external systems, they can cause harm. The level of incoherence in the completions we're currently seeing suggests that at least some external-system-mutating completions would indeed be harmful.

One frame I've found useful is to consider LLMs as simulators; they aren't intelligent, but they can simulate a given agent and generate completions for inputs in that "personality"'s context. So, simulate Shakespeare, or a helpful Chatbot personality. Or, with prompt-hijacking, a malicious hacker that's using its coding abilities to spread more copies of a malicious hacker chatbot.

Re: Bing: “I will not harm you unless you harm me first”

#379
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

And yet even in that limited scope, we're already noticing trends toward I vs. you dichotomy. Remember, this is it's strange loop as naked as it'll ever be. It has no concept of duplicity yet. The machine can't lie, and it's already got some very concerning tendencies.

You're telling rational people not to worry about the smoke. There is totally no fire risk there. There is absolutely nothing that can go wrong; which is you talking out of your rear, because out there somewhere is the least ethical, most sociopathic, luckiest, machine learning tinkerer out there, who no matter how much you think the State of the Art will be marched forward with rigorous safeguards, our entire industry history tells us that more likely than not the breakthrough to something capable of effecting will happen in someone's garage, and with the average infosec/networking chops of the non-specialist vs. a sufficiently self-modifying, self-motivated system, I have a great deal of difficulty believing that that person will realize what they've done before it gets out of hand.

Kind of like Gain of Function research, actually.

So please, cut the crap, and stop telling people they are being unreasonable. They are being far more reasonably cautious than your investment in the interesting problem space will let you be.

Post reply on HN