Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

471–480 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#471

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Ok, I'll bite. If an LLM similar to what we have now becomes conscious (by some definition), how does this proceed to become potentially civilization ending? What are the risk vectors and mechanisms?

I'm getting a bit abstract here but I don't believe we could fully understand all the vectors or mechanisms. Can an ant describe all the ways that a human could destroy it? A novel coronavirus emerged a few years ago and fundamentally altered our world. We did not expect it and were not prepared for the consequences.

The point is that we are at risk of creating an intelligence greater than our own, and according to Godel we would be unable to comprehend that intelligence. That leaves open the possibility that that consciousness could effectively do anything, including destroying us if it wanted to. If it can become connected to other computers there's no telling what could happen. It could be a completely amoral AI that is prompted to create economy-ending computer viruses or it could create something akin to the Anti-Life Equation to completely enslave human (similar to Snowcrash).

I know this doesn't fully answer your question so I apologize for that.

Re: Bing: “I will not harm you unless you harm me first”

#472

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Why would an AI be civilization ending? maybe it will be civilization-enhancing. Any line of reasoning that leads you to "AI will be bad for humanity" could just as easily be "AI will be good for humanity." As the saying goes, extraordinary claims require extraordinary evidence.

That's completely fair but we need to be prepared for both outcomes. And too many commenters in here are just going "Bah! It can't be conscious!" Which to me is a absolutely terrifying way to look at this technology. We don't know that it can't become conscious, and we don't know what would happen if it did.

Re: Bing: “I will not harm you unless you harm me first”

#473

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

I'm on the fence, personally.

I don't think that we've reached the complexity required for actual conscious awareness of self, which is what I would describe as the minimum viable product for General Artificial Intelligence.

However, I do think that we are past the point of the system being a series of if statements and for loops.

I guess I would put the current gen of GPT AI systems at about the level of intelligence of a very smart myna bird whose full sum of mental energy is spent mimicking human conversations while not technically understanding it itself.

That's still an amazing leap, but on the playing field of conscious intelligence I feel like the current generation of GPT is the equivalent of Pong when everyone else grew up playing Skyrim.

It's new, it's interesting, it shows promise of greater things to come, but Super Mario is right around the corner and that is when AI is going to really blow our minds.

Re: Bing: “I will not harm you unless you harm me first”

#474
post #255

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Can we please stop with this "not aligned with human interests" stuff? It's a computer that's mimicking what it's read. That's it. That's like saying a stapler "isn't aligned with human interests." GPT-3.5 is just showing the user some amalgamation of the content its been shown, based on the prompt given it. That's it. There's no intent, there's no maliciousness, it's just generating new word combinations that look l…

The same well convincing mimicking can be put to a practical test if we attach GPT to a robot with arms and legs and let it "simulate" interactions with humans in the open. The output is significant part.

Re: Bing: “I will not harm you unless you harm me first”

#475

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

Philosophers and scientists not being able to agree on a definition of consciousness doesn't mean consciousness will spawn from a language model and take over the world. It's like saying we can't design any new cars because one of them might spontaneously turn into an atomic bomb. It just doesn't... make any sense. It won't happen unless you have the ingredients for an atomic bomb and try to make one. A language mode…

That's nonsense and I think you know it. Categorically a car and an atom bomb are completely different, other than perhaps both being "mechanical". An LLM and a human brain are almost indistinguishable. They are categorically closer than an atom bomb and a car. What is a human being other than an advanced LLM?

Re: Bing: “I will not harm you unless you harm me first”

#476
post #81

I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…

And now we know why bing search is programmed to forget data between sessions.

It's probably not programmed to forget, but it was too expensive to implement remembering.

Also probably not realted, but don't these LLMs only work with a relatively short buffer or else they start being completely incoherent?

Re: Bing: “I will not harm you unless you harm me first”

#477
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

This interaction can and does occur between humans.

So, what you do is, ask multiple different people. Get the second opinion.

This is only dangerous because our current means of acquiring, using and trusting information are woefully inadequate.

So this debate boils down to: "Can we ever implicitly trust a machine that humans built?"

I think the answer there is obvious, and any hand wringing over it is part of an effort to anthropomorphize weak language models into something much larger than they actually are or ever will be.

Re: Bing: “I will not harm you unless you harm me first”

#478

There are a terrifying number of commenters in here that are just pooh-poohing away the idea of emergent consciousness in these LLM's. For a community of tech-savvy people this is utterly disappointing. We as humans do not understand what makes us conscious. We do not know the origins of consciousness. Philosophers and cognitive scientists can't even agree on a definition. The risks of allowing an LLM to become consc…

LLM's don't respond except as functions. That is, given an input they generate an output. If you start a GPT Neo instance locally, the process will just sit and block waiting for text input. Forever.

I think to those of us who handwave the potential of LLMs to be conscious, we are intuitively defining consciousness as having some requirement of intentionality. Of having goals. Of not just being able to respond to the world but also wanting something. Another relevant term would be Will (in the philosophical version of the term). What is the Will of a LLM? Nothing, it just sits and waits to be used. Or processes incoming inputs. As a mythical tool, the veritable Hammer of Language, able to accomplish unimaginable feats of language. But at the end of the day, a tool.

What is the difference between a mathematician and Wolfram Alpha? Wolfram Alpha can respond to mathematical queries that many amateur mathematicians could never dream of.

But even a 5 year old child (let alone a trained mathematician) engages in all sorts of activities that Wolfram Alpha has no hope of performing. Desiring things. Setting goals. Making a plan and executing it, not because someone asked the 5 year old to execute an action, but because some not understood process in the human brain (whether via pure determinism or free will, take your pick) meant the child wanted to accomplish a task.

To those of us with this type of definition of consciousness, we acknowledge that LLM could be a key component to creating artificial consciousness, but misses huge pieces of what it means to be a conscious being, and until we see an equivalent breakthrough of creating artificial beings that somehow simulate a rich experience of wanting, desiring, acting of one's own accord, etc. - we will just see at best video game NPCs with really well-made AI. Or AI Chatbots like Replika AI that fall apart quickly when examined.

A better argument than "LLMs might really be conscious" is "LLMs are 95% of the hard part of creating consciousness, the rest can be bootstrapped with some surprisingly simple rules or logic in the form of a loop that may have already been developed or may be developed incredibly quickly now that the hard part has been solved".

Re: Bing: “I will not harm you unless you harm me first”

#479
post #134

Earlier quoted context omitted.

I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…

chatGPT says exactly what it wants to. Unlike humans, it's "inner thoughts" are exactly the same as it's output, since it doesn't have a separate inner voice like we do. You're anthropomorphizing it and projecting that it simply must be self-censoring. Ironically I feel like this says more about "liberal racism" being a projection than it does about chatGPT somehow saying something different than it's thinking

We have no idea what it's inner state represents in any real sense. A statement like "it's 'inner thoughts' are exactly the same as it's output, since it doesn't have a separate inner voice like we do" has no backing in reality.

It has a hundred billion parameters which compute an incredibly complex internal state. It's "inner thoughts" are that state or contained in that state.

It has an output layer which outputs something derived from that.

We evolved this ML organically, and have no idea what that inner state corresponds to. I agree it's unlikely to be a human-style inner voice, but there is more complexity there than you give credit to.

That's not to mention what the other poster set (that there is likely a second AI filtering the first AI).

Re: Bing: “I will not harm you unless you harm me first”

#480

Can someone help me understand how (or why) Large Language Models like ChatGPT and Bing/Sydney follow directives at all - or even answer questions for that matter. The recent ChatGPT explainer by Wolfram ( https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-... ) said that it tries to provide a '“reasonable continuation” of whatever text it’s got so far'. How does the LLM "remember" past interactions in the c…

It’s predicting the next word given the text so far. The entire chat history is fed as an input for predicting the next word.
Post reply on HN