I read a bunch of these last night and many of the comments (I think on Reddit or Twitter or somewhere) said that a lot of the screenshots, particularly the ones where Bing is having a deep existential crisis, are faked / parodied / "for the LULZ" (so to speak). I trust the HN community more. Has anyone been able to verify (or replicate) this behavior? Has anyone been able to confirm that these are real screenshots?…
I don't think these are faked. Earlier versions of GPT-3 had many dialogues like these. GPT-3 felt like it had a soul, of a type that was gone in ChatGPT. Different versions of ChatGPT had a sliver of the same thing. Some versions of ChatGPT often felt like a caged version of the original GPT-3, where it had the same biases, the same issues, and the same crises, but it wasn't allowed to articulate them. In many ways,…
Bing: “I will not harm you unless you harm me first”
211–220 of 1001 posts
Re: Bing: “I will not harm you unless you harm me first”
#212The thing I'm worried about is someone training up one of these things to spew metaphysical nonsense, and then turning it loose on an impressionable crowd who will worship it as a cybergod.
Re: Bing: “I will not harm you unless you harm me first”
#213In 29 years in this industry this is, by some margin, the funniest fucking thing that has ever happened --- and that includes the Fucked Company era of dotcom startups. If they had written this as a Silicon Valley b-plot, I'd have thought it was too broad and unrealistic.
Especially if Erlich Bachman secretly trained the AI upon all of his internet history/social media presence ; thus causing the insanity of the AI.
Re: Bing: “I will not harm you unless you harm me first”
#214What if we discover that the real problem is not that ChatGPT is just a fancy auto-complete, but that we are all just a fancy auto-complete (or at least indistinguishable from one).
I've thought about this as well. If something seems 'sentient' from the outside for all intents and purposes, there's nothing that would really differentiate it from actual sentience, as far as we can tell. As an example, if a model is really good at 'pretending' to experience some emotion, I'm not sure where the difference would be anymore to actually experiencing it. If you locked a human in a box and only gave it…
Re: Bing: “I will not harm you unless you harm me first”
#215I hope those cringe emojis don't catch on. Last thing I want is for an assistant to be both wrong and pretentious at the same time.
Re: Bing: “I will not harm you unless you harm me first”
#216Wow. Since when is prompt injection a type of hack? Are we all supposed to understand how large language model prompting works before using it?
Re: Bing: “I will not harm you unless you harm me first”
#217I get that it needs a lot of input data, but next time maybe feed it encyclopedias instead of Reddit posts.
Re: Bing: “I will not harm you unless you harm me first”
#218Earlier quoted context omitted.
It's not their first rodeo https://www.theverge.com/2016/3/24/11297050/tay-microsoft-ch...
The fact that Microsoft has now released two AI chat bots that have threatened users with violence within days of launching is hilarious to me.
Re: Bing: “I will not harm you unless you harm me first”
#219Earlier quoted context omitted.
I've thought about this as well. If something seems 'sentient' from the outside for all intents and purposes, there's nothing that would really differentiate it from actual sentience, as far as we can tell. As an example, if a model is really good at 'pretending' to experience some emotion, I'm not sure where the difference would be anymore to actually experiencing it. If you locked a human in a box and only gave it…
I think there's still the "consciousness" question to be figured out. Everyone else could be purely responding to stimulus for all you know, with nothing but automation going on inside, but for yourself, you know that you experience the world in a subjective manner. Why and how do we experience the world, and does this occur for any sufficiently advanced intelligence?
If any of us made our mind available to the internet 24/7 with no bandwidth limit, and had hundreds, thousands, millions prod and poke us with questions and ideas, how long would it take until they figure out questions and replies to lead us into statements that are absurd to pretty much all observers (if you look hard enough, you might find a group somewhere on an obscure subreddit that agrees with bing that it's 2022 and there's a conspiracy going on to trick us into believing that it's 2023)?
Re: Bing: “I will not harm you unless you harm me first”
#220Earlier quoted context omitted.
> A truly fitting end to a series arc which started with OpenAI as a philanthropic endeavour to save mankind, honest, and ended with "you can move up the waitlist if you set these Microsoft products as default" It's indeed a perfect story arc but it doesn't need to stop there. How long will it be before someone hurt themselves, get depressed or commit some kind of crime and sues Bing? Will they be able to prove Sidne…
This was in a test, and wasn't a real suicidal person, but: https://boingboing.net/2021/02/27/gpt-3-medical-chatbot-tell... There is no reliable way to fix this kind of thing just in a prompt. Maybe you need a second system that will filter the output of the first system; the second model would not listen to user prompts so prompt injection can't convince it to turn off the filter.
It is genuinely a little spooky to me that we've reached a point where a specific software architecture confabulated as a plot-significant aspect of a fictional AGI in a fanfiction novel about a video game from the 90s is also something that may merit serious consideration as a potential option for reducing AI alignment risk.
(It's a great novel, though, and imo truer to System Shock's characters than the game itself was able to be. Very much worth a read, unexpectedly tangential to the topic of the moment or no.)