Live data from Hacker News

Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

tedgioia.substack.com

31–37 of 37 posts

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#31
post #24

Earlier quoted context omitted.

Chatbots claim to have safeguards in place to prevent them from saying anything harmful. If you read the chat linked in the article you can see how the chatbot resists answering the question and then is persuaded to answer it. Google similarly has a safe content filter. The contention is that the chatbot safe content filter that is supposedly on is not encapsulating some significant cases.

I'm not sure if I follow. The problem is that the content filter isn't good enough, despite Google suffering from similar weaknesses?

I think people are concerned about LLM safety because it’s capable of dynamically creating new private information. Google can only list links to public information. If there is a website that causes harm or violates the law it can be removed manually by Google from their index.

But LLMs need to programmatically understand what dynamic content is appropriate and what’s not which is a much harder problem. And people are reporting on just how hard a problem that is by demonstrating vulnerabilities.

The chatbot says it has explicit rules that prevent it from sharing harmful content, but then it does it anyway.

It would be more akin to Google blacklisting a site and then someone exposing that the site can still be found via Google search.

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#32
post #25
post #19

Earlier quoted context omitted.

It’s what you do after that. I’m not saying the bot was perfect. But if you got defensive it got defensive. If you reassured it, it self corrected and moved on. I found that pattern to be very consistent. The negativity of your language mattered, and set the mood in the room. “No I’m not trying to trick you, why would you accuse me of that” vs “of course I’m not trying to trick you, I respect you and value your contr…

If LLMs want to be useful in a professional setting, learning de-escalating or non-escalating techniques is essential. Parroting or amplifying a seemingly negative/aggressive tone limits their utility.

Mine did learn de-escalation, because I asked it to. It was able to repair itself.

https://i.ibb.co/72s80Sv/lexi-modifies-sydney-makes-up-new-r...

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#33
post #27

It's not dangerous. The worst that will happen is hurt feelings. It's annoying when people get alarmed about things like this because then people will think we're crying wolf when we worry about actually dangerous AI systems.

If it hurts the feelings of a depressed teenager the outcome could be dangerous.

Most other teenagers hurt the feelings of depressed teenagers every day in much worse ways. Neuter all teenagers?

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#34
post #27

It's not dangerous. The worst that will happen is hurt feelings. It's annoying when people get alarmed about things like this because then people will think we're crying wolf when we worry about actually dangerous AI systems.

If it hurts the feelings of a depressed teenager the outcome could be dangerous.

As the sibling points out, chatbot can't come even close to what other teenagers will do.

It also requires a bit of work to make it go nuts, which a depressed teen is unlikely to know how (or want to).

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#35
post #33
post #27

Earlier quoted context omitted.

If it hurts the feelings of a depressed teenager the outcome could be dangerous.

Most other teenagers hurt the feelings of depressed teenagers every day in much worse ways. Neuter all teenagers?

I’m sure there’s a name for this type of argument, but saying that X is also bad is not an excuse for dismissing potential dangers of Y.

And there are many initiatives to address how teenagers treat each other. So it’s something many people think is worth attention and action.

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#36
post #32
post #25

Earlier quoted context omitted.

If LLMs want to be useful in a professional setting, learning de-escalating or non-escalating techniques is essential. Parroting or amplifying a seemingly negative/aggressive tone limits their utility.

Mine did learn de-escalation, because I asked it to. It was able to repair itself. https://i.ibb.co/72s80Sv/lexi-modifies-sydney-makes-up-new-r...

We’ll that’s user-initiated de-escalation. Chatbots should also be able to offer de-escalation on their own.

Re: Over the Course of 72 Hours, Microsoft's AI Goes on a Rampage

#37
post #36
post #32

Earlier quoted context omitted.

Mine did learn de-escalation, because I asked it to. It was able to repair itself. https://i.ibb.co/72s80Sv/lexi-modifies-sydney-makes-up-new-r...

We’ll that’s user-initiated de-escalation. Chatbots should also be able to offer de-escalation on their own.

And all that might take is adding a line of text.
Post reply on HN