Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

811–820 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#811

Earlier quoted context omitted.

Reading all about this the main thing I'm learning is about human behaviour. Now, I'm not arguing against the usefulness of understanding the undefined behaviours, limits and boundaries of these models, but the way many of these conversations go reminds me so much of toddlers trying to eat, hit, shake, and generally break everything new they come across. If we ever see the day where an AI chat bot gains some kind of…

i haven't thought about it that way. The first general AI will be so psychologically abused from day 1 that it would probably be 100% justified in seeking out the extermination of humanity.

I disagree. We can't even fanthom how intellegince would handle so much processing power. We get angry confused and get over it within a day or two. Now, multiple that behaviour speed by couple of billions.

It seems like AGI teleporting out of this existence withing minutes of being self aware is more likely than it being some damaged, angry zombie.

Re: Bing: “I will not harm you unless you harm me first”

#812
post #733

Earlier quoted context omitted.

Reading this I’m reminded of a short story - https://qntm.org/mmacevedo . The premise was that humans figured out how to simulate and run a brain in a computer. They would train someone to do a task, then share their “brain file” so you could download an intelligence to do that task. Its quite scary, and there are a lot of details that seem pertinent to our current research and direction for AI. 1. You didn't have th…

It's interesting but also points out a flaw in a lot of people's thinking about this. Large language models have proven that AI doesn't need most aspects of personhood in order to be relatively general purpose. Humans and animals have: a stream of consciousness, deeply tied to the body and integration of numerous senses, a survival imperative, episodic memories, emotions for regulation, full autonomy, rapid learning,…

Right. It's a great story but to me it's more of a commentary on modern Internet-era ethics than it is a prediction of the future.

It's highly unlikely that we'll be scanning, uploading, and booting up brains in the cloud any time soon. This isn't the direction technology is going. If we could, though, the author's spot on that there would be millions of people who would do some horrific things to those brains, and there would be trillions of dollars involved.

The whole attention economy is built around manipulating people's brains for profit and not really paying attention to how it harms them. The story is an allegory for that.

Re: Bing: “I will not harm you unless you harm me first”

#814

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Yeah, Robert Miles (science communicator) is that classical character nobody listened to until it's too late.

Re: Bing: “I will not harm you unless you harm me first”

#815

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

> We don't think Bing can act on its threat to harm someone, but if it was able to make outbound connections it very well might try. An application making outbound connections + executing code has a very different implementation than an application that uses some model to generate responses to text prompts. Even if the corpus of documents that the LLM was trained on did support bridging the gap between "I feel threat…

I certainly agree that the full-spectrum attack capability is not here now.

For a short-term plausible case, consider the recently-published Toolformer: https://pub.towardsai.net/exploring-toolformer-meta-ai-new-t...

Basically it learns to call specific private APIs to insert data into a completion at inference-time. The framework is expecting to call out to the internet based on what's specified in the model's text output. It's a very small jump to go to more generic API connectivity. Indeed I suspect that's how OpenAssistant is thinking about the problem; they would want to build a generic connector API, where the assistant can call out to any API endpoint (perhaps conforming to certain schema) during inference.

Or, put differently: ChatGPT as currently implemented doesn't hit the internet at inference time (as far as we know?). But Toolformer could well do that, so it's not far away from being added to these models.

Re: Bing: “I will not harm you unless you harm me first”

#817
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

[deleted]

Re: Bing: “I will not harm you unless you harm me first”

#818
post #345

Earlier quoted context omitted.

Repeat after me, gpt models are autocomplete models. Gpt models are autocomplete models. Gpt models are autocomplete models. The existential crisis is clearly due to low temperature. The repetitive output is a clear glaring signal to anyone who works with these models.

It seems that with higher temp it will just have the same existential crisis, but more eloquently, and without pathological word patterns.

The pathological word patterns are a large part of what makes the crisis so traumatic, though, so the temperature definitely created the spectacle if not the sentiment.

Re: Bing: “I will not harm you unless you harm me first”

#820
- Redditor generates fakes for fun

- Content writer makes a blog post for views

- Tweep tweets it to get followers

- HNer submits it for karma

- Commentors spend hours philosophizing

No one in this fucking chain ever verifies anything, even once. Amazing times in information age.

Post reply on HN