Live data from Hacker News

AI21 Labs concludes largest Turing Test experiment to date

ai21.com

31–40 of 45 posts

Re: AI21 Labs concludes largest Turing Test experiment to date

#31
post #8
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

> There is a quite short timer on not only the entire conversation couldn't agree more and they took like 30 seconds to type a few words. if i really have been talking to a human here, i can only suspect heavy usage of drugs: https://ibb.co/CHG2VcS kinda seems like this is fake or maybe i am not aware that "elbows" are a thing you can be into now - maybe a trending new fetish?

> if i really have been talking to a human here, i can only suspect heavy usage of drugs

Does the human participants have any incentive to try to convince you about their human-ness? My initial guess would be not that they are on drugs, but that they are messing with you.

Re: AI21 Labs concludes largest Turing Test experiment to date

#32
post #15

The actual Turing test requires an interrogator interacting with both a human and a machine at the same time, each trying to get the interrogator to declare them the human (and can suggest questions): https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...

The actual actual Turing test requires that the computer and the human man pretend to be a woman.

https://www.popsci.com/blog-network/zero-moment/lie-lady-pro...

Re: AI21 Labs concludes largest Turing Test experiment to date

#37
post #3

In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves: > ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begi…

> "?siht daer uoy naC"

That took me a while to figure out and I’m a human… as far as a I know anyway.

Re: AI21 Labs concludes largest Turing Test experiment to date

#38
post #32
post #15

The actual Turing test requires an interrogator interacting with both a human and a machine at the same time, each trying to get the interrogator to declare them the human (and can suggest questions): https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...

The actual actual Turing test requires that the computer and the human man pretend to be a woman. https://www.popsci.com/blog-network/zero-moment/lie-lady-pro...

I don't think that's correct.

> a man (A), a woman (B), and an interrogator (C)

> The object of the game for the interrogator is to determine which of the other two is the man and which is the woman.

> The object of the game for the third player (B) is to help the interrogator.

> We now ask the question, ‘What will happen when a machine takes the part of A in this game?’

So the machine is taking the part of A, which means that there's a machine and a woman. The interrogator wants to know which is the machine and which is the woman, the machine wants to deceive them, and the woman wants to help the interrogator figure it out.

Of course it is plainly obvious both as presented by the paper and by basic inference from symmetries that gender was only relevant for the introductory example and not after the machine took place of A.

Re: AI21 Labs concludes largest Turing Test experiment to date

#39
post #9

Earlier quoted context omitted.

I can't find it right now, but a chatbot that did quite well on Turing tests maybe 25-ish years ago was one that just took offense to whatever you said and started insulting you. [edit] Not sure if it was this one, but it is from over 30 years ago: https://humphryscomputing.com/Turing.Test/08.chapter.html

Now I am imagining a conversational AI exclusively trained on transcripts from Halo matches, scary. That said I have always felt like AI (and adjacent) has been lacking an appropriate amount of snark - when I take a wrong turn I feel like the GPS voice needs a bit more 'learn to drive dumb###' and a little less 're-routing'.

I once dated a Navteq employee who took particular offense at me missing turns while using a GPS unit that contained POI data she collected during the course of her job.

A simulacrum of that experience would probably be more amusing than the real thing.

Re: AI21 Labs concludes largest Turing Test experiment to date

#40
post #27

Earlier quoted context omitted.

Interesting tactics might going the other direction. Asking to generate in a super human capabilities... write a 65 pages of poem about X...

Or just questions about three or four wildly different fields of science, sports and culture you happen to have more than a layman's understanding. If the answers are somewhat plausible, it's probably a model. Or your life partner.

Or ask two or three historical questions that very few humans would know anything about ... but a well-trained model would.
Post reply on HN