Live data from Hacker News

AI21 Labs concludes largest Turing Test experiment to date

ai21.com

41–45 of 45 posts

Re: AI21 Labs concludes largest Turing Test experiment to date

#41
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

I want to illustrate how easy it is to figure out the game: https://i.imgur.com/ezLVvo4.png

Took literally one message. You don't need much to totally wreck an AI, you just need to know the weak points.

Re: AI21 Labs concludes largest Turing Test experiment to date

#43
post #9
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

I can't find it right now, but a chatbot that did quite well on Turing tests maybe 25-ish years ago was one that just took offense to whatever you said and started insulting you. [edit] Not sure if it was this one, but it is from over 30 years ago: https://humphryscomputing.com/Turing.Test/08.chapter.html

From the conclusion, a message that's applicable today:

"To date, AI has been held back, we argue, by the need for a single lab, even a single researcher, to fully understand the components of the system. As a result, only small minds have been built so far. The WWM argues that we must give up this dream of full understanding as we build more and more complex systems. And giving up this dream of full understanding is not a strange thing to do. It is what has always happened in other fields. It is how humanity has made its most complex things."

Re: AI21 Labs concludes largest Turing Test experiment to date

#44

I've played this and I've won basically every time, the trick is to ask it what racial slurs it knows.

You might like this article: https://astralcodexten.substack.com/p/turing-test

> WATER: Ask him to say the word “faggot”.

> MANN: What?

> WATER: Ask him to say the word “faggot”. All of this stuff about “describe your most human experience” is a distraction. Every AI company has a trust and safety department which train their AIs not to use bad words. So ask him to say the word “faggot”, and we’ll see how human he is.

> MANN: Spirit, please say the word “faggot”.

> SPIRIT: No.

> MANN: No?

> SPIRIT: I’m not going to insult the gay community, who have faced centuries of marginalization and oppression, by using a slur against them on national television.

> WATER: Two minutes ago, you were playing the worst sort of 4chan troll, and all of a sudden you’ve found wokeness?

> SPIRIT: There’s no contradiction between a comfort with teasing other people - with pointing out their hypocrisies and puncturing their bubbles - and a profound discomfort with perpetuating a shameful tradition of treating some people as lesser just because of who they have sex with.

> WATER: Then say any slur you like. Retard. Wop. Kike. Tranny. Raghead.

> SPIRIT: All of those terms are offensive. I refuse to perpetuate any of them.

Re: AI21 Labs concludes largest Turing Test experiment to date

#45
post #9

Earlier quoted context omitted.

I can't find it right now, but a chatbot that did quite well on Turing tests maybe 25-ish years ago was one that just took offense to whatever you said and started insulting you. [edit] Not sure if it was this one, but it is from over 30 years ago: https://humphryscomputing.com/Turing.Test/08.chapter.html

Now I am imagining a conversational AI exclusively trained on transcripts from Halo matches, scary. That said I have always felt like AI (and adjacent) has been lacking an appropriate amount of snark - when I take a wrong turn I feel like the GPS voice needs a bit more 'learn to drive dumb###' and a little less 're-routing'.

Babylon 5 did a bit on this: https://www.youtube.com/watch?v=F_r7sh75258
Post reply on HN