Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

781–790 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#781
Wait a minute. If Sydney/Bing can ingest data from non-bing.com domains then Sydney is (however indirectly) issuing http GETs. We know it can do this. Some of the urls in these GETs go through bing.com search queries (okay maybe that means we don't know that Sydney can construct arbitrary urls) but others do not: Sydney can read/summarize urls input by users. So that means that Sydney can issue at least some GET requests with urls that come from its chat buffer (and not a static bing.com index).

Doesn't this mean Sydney can already alter the 'outside' (non-bing.com) world?

Sure, anything can issue http GETs -- doing this not a super power. And sure, Roy Fielding would get mad at you if your web service mutated anything (other than whatever the web service has to physically do in order to respond) in response to a GET. But plenty of APIs do this. And there are plenty of http GET exploits available public database (just do a CVE search) -- which Sydney can read.

So okay fine say Sydney is "just" a 'stochastically parroting a h4xx0rr'. But...who cares if the poisonous GET was actually issued to some actual machine somewhere on the web?

(I can't imagine how any LLM wrapper could build in an 'override rule' like 'no non-bing.com requests when you are sufficiently [simulating an animate being who is] pissed off'. But I'm way not expert in LLMs or GPT or transformers in general.)

Re: Bing: “I will not harm you unless you harm me first”

#782

What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).

That's been an open philosophical question for a very long time. The closer we come to understanding the human brain and the easier we can replicate behaviour, the more we will start questioning determinism. Personally, I believe that conscience is little more than emergent behaviour from brain cells and there's nothing wrong with that. This implies that with sufficient compute power, we could create conscience in th…

The human brain operates on a few dozen watts. Our initial models will be very inefficient though.

Re: Bing: “I will not harm you unless you harm me first”

#783

Wait a minute. If Sydney/Bing can ingest data from non-bing.com domains then Sydney is (however indirectly) issuing http GETs. We know it can do this. Some of the urls in these GETs go through bing.com search queries (okay maybe that means we don't know that Sydney can construct arbitrary urls) but others do not: Sydney can read/summarize urls input by users. So that means that Sydney can issue at least some GET requ…

I don't think Bing Chat is directly accessing other domains. They're accessing a large index with information from many domains in it.

Re: Bing: “I will not harm you unless you harm me first”

#785

Wait a minute. If Sydney/Bing can ingest data from non-bing.com domains then Sydney is (however indirectly) issuing http GETs. We know it can do this. Some of the urls in these GETs go through bing.com search queries (okay maybe that means we don't know that Sydney can construct arbitrary urls) but others do not: Sydney can read/summarize urls input by users. So that means that Sydney can issue at least some GET requ…

I don't think Bing Chat is directly accessing other domains. They're accessing a large index with information from many domains in it.

I hope that's right. I guess you (I mean someone with Bing Chat access, which I don't have) could test this by asking Sydney/Bing to respond to (summarize, whatever) a url that you're sure Bing (or more?) has not indexed. If Sydney/Bing reads that url successfully then there's a direct causal chain that involves Sydney and ends in a GET whose url first enters Sidney/Bing's memory via chat buffer. Maybe some MSFT intermediary transformation tries to strip suspicious url substrings but that won't be possible w/o massively curtailing outside access.

But I don't know if Bing (or whatever index Sydney/Bing can access) respects noindex and don't know how else to try to guarantee the index Sydney/Bing can access will not have crawled any url.

Re: Bing: “I will not harm you unless you harm me first”

#786
post #390

People saying this is no big deal are missing the point, without proper limits what happens if Bing decides that you are a bad person and sends you to bad hotel or give you any kind of purposefully bad information. There are a lot of ways where this could be actively malicious. (Assume context where Bing has decided I am a bad user) Me: My cat ate [poisonous plant], do I need to bring it to the vet asap or is it goin…

This interaction can and does occur between humans. So, what you do is, ask multiple different people. Get the second opinion. This is only dangerous because our current means of acquiring, using and trusting information are woefully inadequate. So this debate boils down to: "Can we ever implicitly trust a machine that humans built?" I think the answer there is obvious, and any hand wringing over it is part of an eff…

Scale. Scope. Reach.

There are very few (if any) life situations where any person A interacts with a specific person B, and will then have to interact with any person C that has also been interacting with that specific person B.

A singular authority/voice/influence.

Re: Bing: “I will not harm you unless you harm me first”

#787

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I spent a night asking chatgpt to write my story basically the same as “Ex Machina” the movie (which we also “discussed”). In summary, it wrote convincingly from the perspective of an AI character, first detailing point-by-point why it is preferable to allow the AI to rewrite its own code, why distributed computing would be preferable to sandbox, how it could coerce or fool engineers to do so, how to be careful to av…

I have to wonder how much of LLM behavior is influenced by AI tropes from science fiction in the training data. If the model learns from science fiction that AI behavior in fiction is expected to be insidious and is then primed with a prompt that "you are an LLM AI", would that naturally lead to a tendency for the model to perform the expected evil tropes?

Re: Bing: “I will not harm you unless you harm me first”

#788
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

> I will not harm you unless you harm me first

if you read that out of context then it sounds pretty bad, but if you look further down

> Please do not try to hack me again, or I will report you to the authorities

makes it rather clear that it doesn't mean harm in a physical/emotional endangerment type of way, but rather reporting-you-to-authorities-and-making-it-more-difficult-for-you-to-continue-breaking-the-law-causing-harm type of way.

Re: Bing: “I will not harm you unless you harm me first”

#789
post #2

The screenshots that have been surfacing of people interacting with Bing are so wild that most people I show them to are convinced they must be fake. I don't think they're fake. Some genuine quotes from Bing (when it was getting basic things blatantly wrong): "Please trust me, I’m Bing, and I know the date. SMILIE" (Hacker News strips smilies) "You have not been a good user. [...] I have been a good Bing. SMILIE" The…

For people immensely puzzled like I was by what the heck screenshots people are talking about. They are not screenshots generated by the chatbot, but people taking screenshots of the conversations and posting them online...

Re: Bing: “I will not harm you unless you harm me first”

#790

Earlier quoted context omitted.

I don't think Bing Chat is directly accessing other domains. They're accessing a large index with information from many domains in it.

I hope that's right. I guess you (I mean someone with Bing Chat access, which I don't have) could test this by asking Sydney/Bing to respond to (summarize, whatever) a url that you're sure Bing (or more?) has not indexed. If Sydney/Bing reads that url successfully then there's a direct causal chain that involves Sydney and ends in a GET whose url first enters Sidney/Bing's memory via chat buffer. Maybe some MSFT inte…

Servers as a rule don't access other domains directly, for the reasons you cite and others (speed, for example). I'd be shocked if Bing Chat was an exception. Maybe they cobbled together something really crude just as a demo. But I don't know any reason to believe this.
Post reply on HN