AI21 Labs concludes largest Turing Test experiment to date
21–30 of 45 posts
Re: AI21 Labs concludes largest Turing Test experiment to date
#22Re: AI21 Labs concludes largest Turing Test experiment to date
#23Re: AI21 Labs concludes largest Turing Test experiment to date
#24After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…
> There is a quite short timer on not only the entire conversation couldn't agree more and they took like 30 seconds to type a few words. if i really have been talking to a human here, i can only suspect heavy usage of drugs: https://ibb.co/CHG2VcS kinda seems like this is fake or maybe i am not aware that "elbows" are a thing you can be into now - maybe a trending new fetish?
Re: AI21 Labs concludes largest Turing Test experiment to date
#25I'm most interested in how higher-level strategies will fare in the future -- strategies like talking for a while and seeing if the thing contradicts itself, seeing if it seems to have a good model of yourself as an agent, etc.
Re: AI21 Labs concludes largest Turing Test experiment to date
#26In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves: > ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begi…
This is a pretty astute tactic. Because AI models operate on tokens and not characters, it's pretty easy to confuse them when you go down to the character level. Even asking "how many letters in _" is tough for LLMs.
Re: AI21 Labs concludes largest Turing Test experiment to date
#27In summary, humans win the Turing test ~2/3 of the time against current SOTA LLMs. One of the more interesting tactics used was to target a weakness of the LLMs themselves: > ... participants posed questions that required an awareness of the letters within words. For example, they might have asked their chat partner to spell a word backwards, to identify the third letter in a given word, to provide the word that begi…
Interesting tactics might going the other direction. Asking to generate in a super human capabilities... write a 65 pages of poem about X...
Re: AI21 Labs concludes largest Turing Test experiment to date
#28After playing the game they used (linked at top of article) I find it hard to draw much conclusion from this study. There is a quite short timer on not only the entire conversation, but on each response you can type. When the timer runs out it sends your message in partially written form. It seriously stifles what you can ask the other "person" and it makes responses artificially short even to a deeper question. When…
I can't find it right now, but a chatbot that did quite well on Turing tests maybe 25-ish years ago was one that just took offense to whatever you said and started insulting you. [edit] Not sure if it was this one, but it is from over 30 years ago: https://humphryscomputing.com/Turing.Test/08.chapter.html
That said I have always felt like AI (and adjacent) has been lacking an appropriate amount of snark - when I take a wrong turn I feel like the GPS voice needs a bit more 'learn to drive dumb###' and a little less 're-routing'.
Re: AI21 Labs concludes largest Turing Test experiment to date
#29This AI is rather pathetic. To work within the time limit it would need to make typos and mistakes but even the first reply makes it easy to call out. Not very impressive. https://ibb.co/xL0XpZ7
Re: AI21 Labs concludes largest Turing Test experiment to date
#30This is a bit like testing general relativity using a hand timed stopwatch and an elevator. Sure that is a valid though experiment but the test is nowhere near powerful enough to say anything useful.