Live data from Hacker News

AI21 Labs concludes largest Turing Test experiment to date

ai21.com

1–10 of 45 posts

Re: AI21 Labs concludes largest Turing Test experiment to date

#2
Incredibly, they seem to have used several different LLMs, yet made no distinction between the particular AI models used in the analysis. Amazing that they would not realize there is a huge difference in capabilities.

They also did not seem to consider the different performance of individual prompts.

Re: AI21 Labs concludes largest Turing Test experiment to date

#3
In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves:

> ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begins with a specific letter, or to respond to a message like "?siht daer uoy naC", which can be incomprehensible for an AI model, but a human can easily understand...

Re: AI21 Labs concludes largest Turing Test experiment to date

#4
post #3

In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves: > ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begi…

[deleted]

Re: AI21 Labs concludes largest Turing Test experiment to date

#5
After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When conversation is so stunted of course it is harder to distinguish bot and human.

I'm also curious what study participants were told beforehand. If someone only had experience playing around with ChatGPT they might assume they should use a "detect GPT" strategy. Some of those strategies are pretty specific to the safety features that OpenAI implemented. But the LLM here will gladly curse at you or whatever. On the other hand I suspect it is less good than GPT - not that it matters so much when the entire conversation is exchanging single sentences.

Re: AI21 Labs concludes largest Turing Test experiment to date

#6
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

Additionally, I'd like to know how they corrected for this: "In a creative twist, many people pretended to be AI bots themselves in order to assess the response of their chat partners"

Assuming it actually was "many people", then whenever they have a human conversational partner (who also would be voting at the end), that person is going to have a hard time and skew the results.

Like imagine playing this game as a lay person after having used ChatGPT a little bit and then getting a response to your question that says "as a large language model ...". Depending on how well the game was explained to participants, it's possible that some people even did this intentionally to fuck with results.

In a proper Turing test there is supposed to be 1 bot and 2 humans, where one human is incentivized only to demonstrate they are human and the other human is the one asking probing questions and needing to guess which is which (but is already known to be human).

Anyway I've only read the linked article and played the game a couple times, I didn't look through the original research publication. It's certainly possible they did address some of these issues, but it is such a buzzword topic at the moment that I have my doubts. And regardless the linked article should cover limitations. For exactly this reason it is important that we have higher expectations for the quality of general audience writing about AI.

Re: AI21 Labs concludes largest Turing Test experiment to date

#7
post #6
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

Additionally, I'd like to know how they corrected for this: "In a creative twist, many people pretended to be AI bots themselves in order to assess the response of their chat partners" Assuming it actually was "many people", then whenever they have a human conversational partner (who also would be voting at the end), that person is going to have a hard time and skew the results. Like imagine playing this game as a la…

Ok I have to add one more thing that's funny since I just played a couple more times: if your conversational partner is a human and they exit the window mid-chat, it still lets you vote.

Re: AI21 Labs concludes largest Turing Test experiment to date

#8
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

> There is a quite short timer on not only the entire conversation

couldn't agree more and they took like 30 seconds to type a few words.

if i really have been talking to a human here, i can only suspect heavy usage of drugs: https://ibb.co/CHG2VcS

kinda seems like this is fake or maybe i am not aware that "elbows" are a thing you can be into now - maybe a trending new fetish?

Re: AI21 Labs concludes largest Turing Test experiment to date

#9
post #5

After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…

I can't find it right now, but a chatbot that did quite well on Turing tests maybe 25-ish years ago was one that just took offense to whatever you said and started insulting you.

[edit]

Not sure if it was this one, but it is from over 30 years ago: https://humphryscomputing.com/Turing.Test/08.chapter.html

Re: AI21 Labs concludes largest Turing Test experiment to date

#10
post #3

In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves: > ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begi…

This is a pretty astute tactic. Because AI models operate on tokens and not characters, it's pretty easy to confuse them when you go down to the character level. Even asking "how many letters in _" is tough for LLMs.
Post reply on HN