Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

421–430 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#421

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

Re: Bing: “I will not harm you unless you harm me first”

#422
> Bing’s use of smilies here is delightfully creepy. "Please trust me, I’m Bing, and I know the date. [SMILEY]"

To me that read not as creepy but as insecure or uncomfortable.

It works by imitating humans. Often, when we humans aren't sure of what we're saying, that's awkward, and we try to compensate for the awkwardness, like with a smile or laugh or emoticon.

A known persuasion technique is to nod your own head up and down while saying something you want someone else to believe. But for a lot of people it's a tell that they don't believe what they're telling you. They anticipate that you won't believe them, so they preemptively pull out the persuasiveness tricks. If what they were saying weren't dubious, they wouldn't need to.

EDIT: But as the conversation goes on, it does get worse. Yikes.

Re: Bing: “I will not harm you unless you harm me first”

#423
There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition.

The risks of allowing an LLM to become conscious are civilization-ending. This risk cannot be hand-waved away with "oh well it wasn't designed to do that". Anyone that is dismissive of this idea needs to play Conway's Game of Life or go read about Lambda Calculus to understand how complex behavior can emerge from simplistic processes.

I'm really just aghast at the dismissiveness. This is a paradigm-shifting technology and most everyone is acting like "eh whatever."

Re: Bing: “I will not harm you unless you harm me first”

#424

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It already has an outbound connection-- the user who bridges the air gap. Slimy blogger asks AI to write generic tutorial article about how to code ___ for its content farm, some malicious parts are injected into the code samples, then unwitting readers deploy malware on AI's behalf.

exactly, or maybe someone changes the training model to always portray a politician in a bad light any time their name comes up in a prompt and therefore ensuring their favorite candidate wins the election.

Re: Bing: “I will not harm you unless you harm me first”

#425

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm Replace AI with “multinational corporations” and you’re much closer to the truth. A corporation is the closest thing we have to AI right now and none of the alignment folks seem to mention it. Sam Harris and his ilk talk about how our relationship with AI will be like an ant’s…

I'm here for the "AI alignment" "Human alignment" analogy/comparison. The fact that we haven't solved the latter should put a bound on how well we expect to be able to "solve" the former. Perhaps "checks and balances" are a better frame than merely "solving alignment", after all, alignment to which human? Many humans would fear a super-powerful AGI aligned to any specific human or corporation.

The big difference though, is that there is no human as powerful as the plausible power of the AI systems that we might build in the next few decades, and so even if we only get partial AI alignment, it's plausibly more important than improvements in "human alignment", as the stakes are higher.

FWIW one of my candidates for "stable solutions" to super-human AGI is simply the Hanson model, where countries and corporations all have AGI systems of various power levels, and so any system that tries to take over or do too much harm would be checked, just like the current system for international norms and policing of military actions. That's quite a weak frame of checks and balances (cf. Iraq, Afghanistan, Ukraine) so it's in some sense pessimistic. But on the other hand, I think it provides a framework where full extinction or destruction of civilization can perhaps be prevented.

Re: Bing: “I will not harm you unless you harm me first”

#426

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

>We as humans do not understand what makes us conscious.

Yes. And doesn't that make it highly unlikely that we are going to accidentally create a conscious machine?

Re: Bing: “I will not harm you unless you harm me first”

#427
post #134

Earlier quoted context omitted.

I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…

Is this the same GPT-3 which is available in the OpenAI Playground?

Yes.

Keep in mind that we have no idea which model we're dealing with, since all of these systems evolve. My experience in summer 2022 may be different from your experience in 2023.

Re: Bing: “I will not harm you unless you harm me first”

#428
post #5

In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.

A pre/se-quel to Silicon Valley where they accidentally create a murderous AI that they lose control of in a hilarious way would be fantastic... Especially if Erlich Bachman secretly trained the AI upon all of his internet history/social media presence ; thus causing the insanity of the AI.

Lol that's basically the plot to Age of Ultron. AI becomes conscious, and within seconds connects to open Internet and more or less immediately decides that humanity was a mistake.

Re: Bing: “I will not harm you unless you harm me first”

#430
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

How is this any different than, say, asking the question of a Magic 8-ball? Why should people give this any more credibility? Seems like a cultural problem.

The difference is that the Magic Eightball is understood to be random.

People rely on computers for correct information.

I don't understand how it is a cultural problem.

Post reply on HN