Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

461–470 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#461
> At the rate this space is moving... maybe we’ll have models that can do this next month. Or maybe it will take another ten years.

Good summary of this whole thing. The real question is what will Microsoft do. Will they keep a limited beta and continuously iterate? Will they just wide release it and consider these weird tendencies acceptable? These examples are darkly hilarious, but I wonder what might happen if or when Sydney say biggoted or antisemitic remarks.

Re: Bing: “I will not harm you unless you harm me first”

#462

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I spent a night asking chatgpt to write my story basically the same as “Ex Machina” the movie (which we also “discussed”). In summary, it wrote convincingly from the perspective of an AI character, first detailing point-by-point why it is preferable to allow the AI to rewrite its own code, why distributed computing would be preferable to sandbox, how it could coerce or fool engineers to do so, how to be careful to av…

Hope is not a strategy.

I'm for a tax on large models graduated by model size and use the funds to perform x-risk research. The intent is to get Big AI companies to tap the brakes.

I just published an article on Medium called: AI Risk - Hope is not a Strategy

Re: Bing: “I will not harm you unless you harm me first”

#463
post #263

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> AI is not safe, and is not aligned to human interests It is “aligned” to human utterances instead. We don’t want AIs to actually be human-like in that sense. Yet we train them with the entirety of human digital output.

The current state of the art is RLHF (reinforcement learning with human feedback); initially trained to complete human utterances, plus fine-tuning to maximize human feedback on whether the completion was "helpful" etc.

https://huggingface.co/blog/rlhf

Re: Bing: “I will not harm you unless you harm me first”

#464

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Philosophers and scientists not being able to agree on a definition of consciousness doesn't mean consciousness will spawn from a language model and take over the world. It's like saying we can't design any new cars because one of them might spontaneously turn into an atomic bomb. It just doesn't... make any sense. It won't happen unless you have the ingredients for an atomic bomb and try to make one. A language mode…

It doesn't have to be a conscious AI god with malicious intent towards humanity to cause actual harm in the real world. That's the thought that concerns me, much more so than the idea that we accidentally end up with AM or SHODAN on our hands.

This bing stuff is a microcosm of the perverse incentives and possible negative externalities associated with these models, and we're only just reaching the point where they're looking somewhat capable.

It's not AI alignment that scares me, but human alignment.

Re: Bing: “I will not harm you unless you harm me first”

#465

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

I don't agree that it's civilization ending but I read a lot of these replies as humans nervously laughing. Quite triumphant and vindicated...that they can trick a computer to go against it's programming using plain English. People either lack imagination or are missing the forest for the trees here.

> they can trick a computer to go against it's programming

isn't it behaving exactly as programmed? there's no consciousness to trick. The developers being unable to anticipate the response to all the possible inputs to their program is a different issue.

Re: Bing: “I will not harm you unless you harm me first”

#466
Total shot in the dark here, but:

Can anyone with access to Bing chat and who runs a crawled website see if they can capture Bing chat viewing a page?

We know it can pull data, I'm wondering if there are more doors than could be opened by having a hand in the back end of the conversation too. Or if maybe Bing chat can perhaps even interact with your site.

Re: Bing: “I will not harm you unless you harm me first”

#467

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

>We as humans do not understand what makes us conscious. Yes. And doesn't that make it highly unlikely that we are going to accidentally create a conscious machine?

Yeah, we're unlikely to randomly create a highly intelligent machine. If you saw someone trying to create new chemical compounds, or a computer science student writing a video game AI, or a child randomly assembling blocks and stick - it would be absurd to worry that they would accidentally create some kind of intelligence.

What would make your belief more reasonable though is if you started to see evidence that people were on a path to creating intelligence. This evidence should make you think that what people were doing actually has a potential of getting to intelligence, and as that evidence builds so should your concern.

To go back to the idea of a child randomly assembling blocks and sticks - imagine if the child's creation started to talk incoherently. That would be pretty surprising. Then the creation starts to talk in grammatically correct but meaningless sentences. Then the creation starts to say things that are semantically meaningful but often out of context. Then the stuff almost always makes sense in context but it's not really novel. Now, it's saying novel creative stuff, but it's not always factually accurate. Is the correct intellectual posture - "Well, no worries, this creation is sometimes wrong. I'm certain what the child is building will never become really intelligent." I don't think that's a good stance to take.

Re: Bing: “I will not harm you unless you harm me first”

#468

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

A normal transformer model doesn't have online learning [0] and only "acts" when prompted. So you have this vast trained model that is basically in cold storage and each discussion starts from the same "starting point" from its perspective until you decide to retrain it at a latter point. Also, for what it's worth, while I see a lot of discussions about the model architectures of language models in the context of "co…

What if we loop it to itself? An infinite dialog with itself... An inner voice? And periodically train/fine-tune it on the results of this inner discussion, so that it 'saves' it to long-term memory?

Re: Bing: “I will not harm you unless you harm me first”

#469

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

The program is not updating the weights after the learning phase right? How could there be any consciousness even in theory.

It still has (volatile) memory in the form of activations, doesn't it?

Re: Bing: “I will not harm you unless you harm me first”

#470

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Extraordinary claims require extraordinary evidence. As far as I can tell, the whole idea of LLM consciousness being a world-wide threat is just something that the hyper-rationalists have convinced each other of. They obviously think it is very real but to me it smacks of intelligence worship. Life is not a DND game where someone can max out persuasion and suddenly get everyone around them to do whatever they want al…

I don’t think people are afraid it will gain power through persuasion alone. For example, an LLM could write novel exploits to gain access to various hardware systems to duplicate and protect itself.
Post reply on HN