Live data from Hacker News

Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

exploringpossibilityspace.blogspot.com

41–50 of 57 posts

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#42

If I were developing a chatbot, I wouldn't provide it with a blank slate and leave it to the mercy of 4chan users. I'd give it a personality and a set of values and a rudimentary understanding of how the world works.

> I'd give it a personality and a set of values and a rudimentary understanding of how the world works. I don't think you have any idea how complicated such a task would be.

I do have an idea. I first got involved in AI in the early 1980s, and worked on it on-and-off (mostly on) well into the 1990s. Some of the non-AI stuff I did back then is now, for reasons I can't quote fathom, classified as AI.

A certain amount of common-sense, and basic political knowledge, should have been included in the Tay. I don't think there's any way of avoiding it. I'm skeptical of general application of the recent "free lunch" approach to AI, and for higher-level AI prefer Doug Lenat's approach with Cyc.

Alternatively, if they wanted to avoid the effort of doing that, they could have just put up a souped-up Eliza variant. That wouldn't have impressed many people, but neither did Tay, and it wouldn't have offended anyone.

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#43
post #27

I am not understanding the spin that is being displayed in general with regards to Tay. Microsoft created a chat bot, the chat bot chatted. It wasnt a failure on any technical level as far as i have seen.

Do you really not understand that there are failures other than "the chat bot chatted so its ok"? Do you not realise there might be other requirements for software than "it kind of appears to do roughly what someone asked for"?

I understand completely but what i am saying is... they set out as far as i know to create a social media chat bot that sounds like a teenage girl who is immersed in her social media environment. it didnt fail this in any way as far as i can tell. its just that...

a teenage girl who speaks the language of her social media environment is going to say a lot of dumb outlandish stuff. even if hypothetically this were some deeper than NLP true AI breakthrough it was always going to say outlandish crazy offensive things because that's the influence that's feeding into it.

its similar to how to a degree that recent bipedal robot that was in the videos got some backlash because its human like motion was offputting... the robot didnt fail in anyway...

its more a question of why would you spend all this money rolling out these experiments if you didnt want the very forseeable output. its a no brainer that a chat bot that learns from social media is going to say fucked up shit.

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#44
post #23

It doesn't sound like the author's done much root cause analysis. QA typically doesn't cause defects (it can only prevent them), so there's almost always a deeper cause. And, QA's not necessarily the best way to find bugs like this. In this case, it sounds like to me like it was most likely a requirements problem -- a missing requirement for being 4chan-resilient, if you will. A different way to look at it is that it…

I just wrote a follow-up post giving more evidence: http://exploringpossibilityspace.blogspot.com/2016/03/micros...

I call it "poor software QA" because, generally, the software QA process is supposed to detect and prevent defects from 1) being introduced in the first place; and 2) from being propagated into "production" versions. As my most recent post shows, the "repeat after me" rule was a legacy of a software library (ALICE) they used to implement rule-based behavior. Some sort of QA process should have been done on the rule set they reused and modified.

Also, when I say "QA" I am not referring only to people with QA in their job title. I'm referring to the process.

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#45
OP here. I just added another blog post with additional evidence that this was a QA problem, specifically related to an open source library (ALICE) and the AIML rules they reused and modified.

http://exploringpossibilityspace.blogspot.com/2016/03/micros...

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#46

It's arrogant to think that any kind of QA could prevent how an AI would turn out. That's probably why the terminators took over humans in the movie--because humans were arrogant enough to think that doing enough QA would prevent everything. What should happen instead is you need to be humble and assume that things won't go the way you designed them to, that's the "safe" way to build AIs in the long term. The OP uses…

See: http://exploringpossibilityspace.blogspot.com/2016/03/micros...

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#47

You can have some QA, no QA, or even great QA, but the output is purely a result of its environment. I have a chat bot that went casually racist about a day or two after activating it. After looking through the logs, I found a particularly vitriolic person that was responsible for the source of this bot's newfound hatred of Asians. My bot didn't get fixated on one particular topic, it just spewed racism and vitriol f…

See: http://exploringpossibilityspace.blogspot.com/2016/03/micros...

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#48

How do we know that QA didn't report the issues? QA reporting a bug and that bug being fixed are two different things right?

I would distinguish between "QA as a process" and "QA as a role". We don't know that people in the Quality Assurance role failed; we can observe that the process intended to assure quality did.

OP here. Agreed. I'll modify my post to clarify that I'm talking about "QA as a process", which usually involves people with "QA as a role", but not always.

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#49

Earlier quoted context omitted.

> I'd give it a personality and a set of values and a rudimentary understanding of how the world works. I don't think you have any idea how complicated such a task would be.

I do have an idea. I first got involved in AI in the early 1980s, and worked on it on-and-off (mostly on) well into the 1990s. Some of the non-AI stuff I did back then is now, for reasons I can't quote fathom, classified as AI. A certain amount of common-sense, and basic political knowledge, should have been included in the Tay. I don't think there's any way of avoiding it. I'm skeptical of general application of the…

If you really are experienced as you say you are in AI, you wouldn't talk about these things as if they were easy to implement. I won't assume things but I can say I'm probably not so less experienced than yourself, and I know basically everything in this field belongs to "easier said than done" category. Most researchers just run experiments in closed environments just like you said and that's what makes them useless and out of touch with reality. For this reason I actually applaud MS guys for having the guts to do this in public. It's much better than them coming out with some lame, controlled environment "AI" which does exactly what its creators intended. It's not a "failure". It's a learning process. It's not like this chatbot went and killed anyone. Everyone knew it was a robot when they were engaging with it, which is not so different from watching a standup comedian making a racist joke on stage.

Re: Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)

#50
post #33

Earlier quoted context omitted.

That's not as easy as it seems. See: https://en.wikipedia.org/wiki/Scunthorpe_problem and http://stackoverflow.com/questions/273516/how-do-you-impleme...

Oh, thanks for the references. I'm constantly astonished at how people can think "filter bad words" would fix everything. In some instances of Tay tweets, there are comparisons between black men and monkeys. What words should you forbid? Black, man, or monkey? Should the bot be allowed to speak about people of color visiting a zoo?

This is supposed to be AI. I'd expect it to know what words and phrases mean, and how sentences are structured. I'd expect it to know that some people are black, and people can visit zoos, and zoos often contain monkeys. I'd also expect it to know that comparing people to other animals is potentially offensive, and to avoid saying those things just to be on the safe side.
Post reply on HN