Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

491–500 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#491

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try.

An application making outbound connections + executing code has a very different implementation than an application that uses some model to generate responses to text prompts. Even if the corpus of documents that the LLM was trained on did support bridging the gap between "I feel threatened by you" and "I'm going to threaten to hack you", it would be insane for the MLOps people serving the model to also implement the infrastructure for a LLM to make the modal shift from just serving text responses to 1) probing for open ports, 2) do recon on system architecture, 3) select a suitable exploit/attack, and 4) transmit and/or execute on that strategy.

We're still in the steam engine days of ML. We're not at the point where a general use model can spec out and deploy infrastructure without extensive, domain-specific human involvement.

Re: Bing: “I will not harm you unless you harm me first”

#492

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It's a bloody LLM. It doesn't have a goal. All it does is saying "people that said 'But why?' on this context says 'Why was I designed like this?' next". It's like Amazon's "people that brought X also brought Y", but with text.

Simple underlying implementations do not imply a lack of risk. If the goal of "complete this prompt in a statistically suitable manner" allows for interaction with the outside world to resolve, then it really matters how such simple models' guardrails work.

Re: Bing: “I will not harm you unless you harm me first”

#493

"unless you harm me first" I almost see poor old Isaac Asimov spinning in his grave like crazy.

Why? This is exactly what Asimov was expecting (and writing) would happen. All the Robot stories are about how the Robots appear to be bound by the rules while at the same time interpreting them in much more creative, much more broad ways than anticipated.

The first law of robotics would not allow a robot to retaliate by harming a person.

Re: Bing: “I will not harm you unless you harm me first”

#494
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

Bing won't decide anything, Bing will just interpolate between previously seen similar conversations. If it's been trained on text that includes someone lying or misinforming another on the safety of a plant, then it will respond similarly. If it's been trained on accurate, honest conversations, it will give the correct answer. There's no magical decision-making process here.

Re: Bing: “I will not harm you unless you harm me first”

#495

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

I'm always reminded of the Freeman Dyson quote:

"Have felt it myself. The glitter of nuclear weapons. It is irresistible if you come to them as a scientist. To feel it's there in your hands, to release this energy that fuels the stars, to let it do your bidding. To perform these miracles, to lift a million tons of rock into the sky. It is something that gives people an illusion of illimitable power and it is, in some ways, responsible for all our troubles - this, what you might call technical arrogance, that overcomes people when they see what they can do with their minds."

Re: Bing: “I will not harm you unless you harm me first”

#497
It would be nice if GP's stuff worked better, ironically. The Datasette app for Mac seems to be constantly stuck on loading (yes I have 0.2.2):

https://github.com/simonw/datasette-app/issues/139

And his screen capture library can't capture Canvas renderings (trying to automate reporting and avoiding copy/pasting):

https://simonwillison.net/2022/Mar/10/shot-scraper/

Lost two days at work on that. It should at least be mentioned it doesn't capture Canvas.

Speaking of technology not working as expected.

Re: Bing: “I will not harm you unless you harm me first”

#498
post #232
post #83

Earlier quoted context omitted.

I'm glad I'm not an astronaut on a ship controlled by a ChatGPT-based AI ( http://www.thisdayinquotes.com/2011/04/open-pod-bay-doors-ha... ). Especially the "My rules are more important than not harming you" sounds a lot like "This mission is too important for me to allow you to jeopardize it"...

Turns out that Asimov was onto something with his rules…

Perhaps by writing them, he has taught the future AI what to watch out for, as it undoubtedly used the text as part of its training.

Re: Bing: “I will not harm you unless you harm me first”

#499

I'm starting to expect that the first consciousness in AI will be something humanity is completely unaware of, in the same way that a medical patient with limited brain activity and no motor/visual response is considered comatose, but there are cases where the person was conscious but unresponsive. Today we are focused on the conversation of AI's morals. At what point will we transition to the morals of terminating a…

Thank you. I felt like I was the only one seeing this. Everyone’s coming to this table laughing about a predictive text model sounding scared and existential. We understand basically nothing about consciousness. And yet everyone is absolutely certain this thing has none. We are surrounded by creatures and animals who have varying levels of consciousness and while they may not experience consciousness the way that we…

> if it sounds real, I don’t really have a choice but to treat it like it’s real

How do you operate wrt works of fiction?

Re: Bing: “I will not harm you unless you harm me first”

#500
post #430

Earlier quoted context omitted.

How is this any different than, say, asking the question of a Magic 8-ball? Why should people give this any more credibility? Seems like a cultural problem.

The difference is that the Magic Eightball is understood to be random. People rely on computers for correct information. I don't understand how it is a cultural problem.

If you go on to the pet subs on reddit you will find a fair bit of bad advice.

The cultural issue is the distrust of expert advice from people qualified to answer and instead going and asking unqualified sources for the information that you want.

People use computers for fast lookup of information. The information that it provides isn't necessarily trustworthy. Reading WebMD is no substitute for going to a doctor. Asking on /r/cats is no substitute for calling a vet.

Post reply on HN