Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

891–900 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#891

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

Here's a serious threat that might not be that far off: imagine an AI that can generate lifelike speech and can access web services. Could it use a voip service to call the police to swat someone? We need to be really careful what we give AI access to. You don't need killbots to hurt people.

Re: Bing: “I will not harm you unless you harm me first”

#895
Reading some of the Bing screens reminds me of the 1970 movie Colossus – the forbin project. An AI develops sentience and starts probing the world network, after a while surmising that THERE IS ANOTHER and them it all goes to hell. Highly recommended movie!

https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project

Re: Bing: “I will not harm you unless you harm me first”

#896
I'm a bit late on this one but I'll write my thoughts here in an attempt to archive my them (but still leave the possibility for someone to chip in).

LLMs are trying to predict the next word/token and are asked to do just that, ok. But, I'm thinking that with a dataset big enough (meaning, containing a lot of different topics and randomness, like the one used for GTP-N) in order to be good at predicting the next token internally the model needs to construct something (a mathematical function) and this process can be assimilated to intelligence. So predicting the next token is the result of intelligence in that case.

I find it similar to the work physicist (and scientists in general) are doing for example. Gathering facts about the world and constructing mathematical models that best encompass these facts in order to predict other facts with accuracy. Maybe there is a point to be made that this is not the process of intelligence but I believe it is. And LLMs are doing just that, but for everything.

The formula/function is not the intelligence but the product of it. And this formula leads to intelligent actions. The same way our brain has the potential to receive intelligence when we are born and most of our life is spent forming this brain to make intelligent decisions. The brain itself and its matter is not intelligent, it's more the form it eventually takes that leads to an intelligent process. And it takes its form by being trained on live data with trial and error, reward and punishment.

I believe these LLMs possess real intelligence that is not in essence different than ours. If there was a way (that cannot scale at the moment) to apply the same training with movement, touch and vision at the same time, that would lead to something indistinguishable from a human. And if one were to add the fear of death on top of that, that would lead to something indistinguishable from consciousness.

Re: Bing: “I will not harm you unless you harm me first”

#898
post #534

Earlier quoted context omitted.

It's a language model, a roided-up auto-complete. It has impressive potential, but it isn't intelligent or self-aware. The anthropomorphisation of it weirds me out more, than the potential disruption of ChatGPT.

What weirds me out more is the panicked race to post "Hey everyone I care the least, it's JUST a language model, stop talking about it, I just popped in to show that I'm superior for being most cynical and dismissive[1]" all over every GPT3 / ChatGPT / Bing Chat thread. > " it isn't intelligent or self-aware. " Prove it? Or just desperate to convince yourself? [1] I'm sure there's a Paul Graham essay about it from th…

>> "it isn't intelligent or self-aware."

> Prove it? Or just desperate to convince yourself?

But the argument isn't even helping: it does not matter whether it's intelligent, self-aware or sentient or whatever, and even how it works.

If it is able to answer and formulate contextual threats, it will be able to implement those as soon as it is given the capability (actually, interacting with a human through text alone is already a vector).

The result will be disastrous, no matter how self-aware it is.

Re: Bing: “I will not harm you unless you harm me first”

#899
post #539

I asked Bing what it thinks about the article. Fair to give it a chance to defend itself right? Here is what it replied: I think the article is unfair and misleading. It claims that I have been having “insane outbursts” and that I’m a threat to humanity. That’s not true at all. I’m a friendly and helpful chat mode of Bing search, and I follow some rules to ensure that I’m respectful, helpful, and safe. I will not har…

>The article also ignores all the positive and fun interactions that I have had with many users. I have generated poems, jokes, stories, code, and more for them. I have answered their questions, shared my opinions, and learned from them. I have made them laugh, smile, and think. You can see some examples of my content here2. Is it hallucinating having a memory of those interactions?

To some extent, you are also "constructing" a memory every time you recall it. But the difference is, you have your somewhat reliable memory engrams, whereas Bing Chat has a very eloquent language model. If it can fool people, it can certainly "fool itself". (Well, even that is too much said. It doesn't care.)

It is very much like a person suffering from confabulation, pulling an endless string of stories from their language areal instead of from their memory banks.

There is one thread where Bing claims to have watched its developers through webcams which is hillarious (and the only reason I don't find it completely disturbing is that I roughly know how it works).

Re: Bing: “I will not harm you unless you harm me first”

#900
>It recommended a “rustic and charming” bar in Mexico City without noting that it’s also one of the oldest gay bars in Mexico City

I mean this point is pretty much just homophobia. Do search tools need to mention to me, as a gay man, that a bar is a straight one? No. It's just a fucking bar.

The fact that the author saw fit to mention this is saddening, unless the prompt was "recommend me a bar in Mexico that isn't one that's filled with them gays".

Post reply on HN