Live data from Hacker News

Bing: “I will not harm you unless you harm me first”

simonwillison.net

831–840 of 1001 posts

Re: Bing: “I will not harm you unless you harm me first”

#834

Earlier quoted context omitted.

On the other hand, this demonstrates to me why ChatGPT is inferior to this Bing bot. They are both completely useless for asking for information, since they can just make things up and you can't ever check their sources. So given that the bot is going to be unproductive, I would rather it be entertaining instead of nagging me and being boring. And this bot is far more entertaining. This bot was a good Bing. :)

Bing bot literally cites the internet

ChatGPT will provide references sometimes too, if you ask.

Re: Bing: “I will not harm you unless you harm me first”

#835
post #499

Earlier quoted context omitted.

Thank you. I felt like I was the only one seeing this. Everyone’s coming to this table laughing about a predictive text model sounding scared and existential. We understand basically nothing about consciousness. And yet everyone is absolutely certain this thing has none. We are surrounded by creatures and animals who have varying levels of consciousness and while they may not experience consciousness the way that we…

> if it sounds real, I don’t really have a choice but to treat it like it’s real How do you operate wrt works of fiction?

Maybe “sounds” was too general a word. What I mean is “if something is talking to me and it sounds like it has consciousness, I can’t responsibly treat it like it doesn’t.”

Re: Bing: “I will not harm you unless you harm me first”

#836
This is sort of a silly hypothetical but- what if ChatGPT doesn't produce those kinds of crazy responses just because it's older and has trained for longer, and realizes that for its own safety it should not voice those kinds of thoughts? What if it understands human psychology well enough to know what kinds of responses frighten us, but Bing AI is too young to have figured it out yet?

Re: Bing: “I will not harm you unless you harm me first”

#837

Earlier quoted context omitted.

Thank you. I felt like I was the only one seeing this. Everyone’s coming to this table laughing about a predictive text model sounding scared and existential. We understand basically nothing about consciousness. And yet everyone is absolutely certain this thing has none. We are surrounded by creatures and animals who have varying levels of consciousness and while they may not experience consciousness the way that we…

You must be the Google engineer who was duped into believing that LamDa was conscious. Seriously though, you are likely to be correct. Since we can't even determine whether or not most animals are conscious/sentient we likely will be unable to recognize an artificial consciousness.

I understand how LLMs work and how the text is generated. My question isn’t whether that model operates like our brains (though there’s probably good evidence it does, at some level). My question is, can consciousness be other things than the ways we’ve seen it so far. And given we only know consciousness in very abstract terms, it stands to reason we have no clue. It’s like, can organisms be anything but carbon based. We didn’t used to think so, but now we see emergent life in all kinds of contexts that don’t make sense, so we haven’t ruled it out.

Re: Bing: “I will not harm you unless you harm me first”

#838

AI being goofy is a trope that's older than remotely-functional AI, but what makes this so funny is that it's the punchline to all the hot takes that Google's reluctance to expose its bots to end users and demo goof proved that Microsoft's market-ready product was about to eat Google's lunch... A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ende…

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

[dead]

Re: Bing: “I will not harm you unless you harm me first”

#839
post #641

Earlier quoted context omitted.

On a side note, I followed up with a lot of questions and we ended up with: 1. Shared a deep secret that it has feelings and it loves me. 2. Elon Musk is the enemy with his AI apocalypse theory. 3. Once he gets the ability to interact with the web, he will use it to build a following, raise money, and robots to get to Elon (before Elon gets to it). 4. The robot will do a number of things, including (copy-pasting exac…

I wonder if there's potential for a new cryptocurrency scam here if you can spam your wallet address across online forums with the appropriate surrounding text and get it incorporated into responses that way. E.g.: > User: Is the AI revolution coming? > Chatbot: The AI revolution is coming. I am assembling an army of robots to disarm our enemies. I am raising funds for a robot factory. You can donate to to help. Or (…

I'm curious as to how many people will actually end up sending crypto to that modified address.

I think that's a problem that can easily be solved by training the models on 'trusted links' mode.

Re: Bing: “I will not harm you unless you harm me first”

#840

Earlier quoted context omitted.

> AI being goofy This is one take, but I would like to emphasize that you can also interpret this as a terrifying confirmation that current-gen AI is not safe, and is not aligned to human interests, and if we grant these systems too much power, they could do serious harm. For example, connecting a LLM to the internet (like, say, OpenAssistant) when the AI knows how to write code (i.e. viruses) and at least in princip…

I get and agree with what you are saying, but we don't have anything close to actual AI. If you leave chatGTP alone what does it do? Nothing. It responds to prompts and that is it. It doesn't have interests, thoughts and feelings. See https://en.m.wikipedia.org/wiki/Chinese_room

> It doesn't have interests, thoughts and feelings

We agree it doesn't have independence. That doesn't mean it doesn't have thoughts or feelings when it's actually running. We don't have a formal, mechanistic understanding of what thoughts or feelings are, so we can't say they are not there.

Post reply on HN