Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

431–440 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#431

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

Our brains are highly recursive, a feature that deep learning models almost never have any, and that GPU have a great deal of trouble to run in any large amount.

That means that no, we think nothing like those AIs.

Re: Bing: “I will not harm you unless you harm me first”

#432

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

The program is not updating the weights after the learning phase right? How could there be any consciousness even in theory.

Re: Bing: “I will not harm you unless you harm me first”

#433
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

And its output is more or less as aligned with human interests as humans are. I think that's the more frightening point.

Re: Bing: “I will not harm you unless you harm me first”

#434
post #257

For anyone wondering if they're fake: I'm extremely familiar with these systems and believe most of these screenshots. We found similar flaws with automated red teaming in our 2022 paper: https://arxiv.org/abs/2202.03286

I wonder if Google is gonna be ultra conservative with Bard, betting that this ultimately blows up in Microsoft's face.

Re: Bing: “I will not harm you unless you harm me first”

#435

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Philosophers and scientists not being able to agree on a definition of consciousness doesn't mean consciousness will spawn from a language model and take over the world.

It's like saying we can't design any new cars because one of them might spontaneously turn into an atomic bomb. It just doesn't... make any sense. It won't happen unless you have the ingredients for an atomic bomb and try to make one. A language model won't turn into a sentient being that becomes hostile to humanity because it's a language model.

Re: Bing: “I will not harm you unless you harm me first”

#436

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

It just highlights even more the need to understand these systems and test them rigorously before deploying them en-masse

Re: Bing: “I will not harm you unless you harm me first”

#437

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

I don't agree that it's civilization ending but I read a lot of these replies as humans nervously laughing. Quite triumphant and vindicated...that they can trick a computer to go against it's programming using plain English. People either lack imagination or are missing the forest for the trees here.

Re: Bing: “I will not harm you unless you harm me first”

#438

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

A normal transformer model doesn't have online learning [0] and only "acts" when prompted. So you have this vast trained model that is basically in cold storage and each discussion starts from the same "starting point" from its perspective until you decide to retrain it at a latter point.

Also, for what it's worth, while I see a lot of discussions about the model architectures of language models in the context of "consciousness" I rarely see a discussion about the algorithms used during the inference step, beam search, top-k sampling, nucleus sampling and so on are incredibly "dumb" algorithms compared to the complexity that is hidden in the rest of the model.

https://www.qwak.com/post/online-vs-offline-machine-learning...

Re: Bing: “I will not harm you unless you harm me first”

#439

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

>We as humans do not understand what makes us conscious. Yes. And doesn't that make it highly unlikely that we are going to accidentally create a conscious machine?

More likely it just means in the process of doing so that we’re unlikely to understand it.

Re: Bing: “I will not harm you unless you harm me first”

#440

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

It already has an outbound connection-- the user who bridges the air gap. Slimy blogger asks AI to write generic tutorial article about how to code ___ for its content farm, some malicious parts are injected into the code samples, then unwitting readers deploy malware on AI's behalf.

Hush! It's listening. You're giving it dangerous ideas!
Post reply on HN