Live data from Hacker News

Killed by LLM

r0bk.github.io

81–90 of 102 posts

Re: Killed by LLM

#81
post #26

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

From TFA: >GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%).

What article? The Turing box on the page has no links to anything and the entire page is just a bunch of these boxes. Sure it has the link icon in the two places like all the others. But the Turing one has no actual links on them. Even the little info icon that works on the GSM8K does nothing for the Turing one.

Re: Killed by LLM

#82

Earlier quoted context omitted.

That's not the original Turing test either. The original imitation game as proposed by Turing involves reading a text transcript of a human and a computer and having the evaluator determine which is which. The evaluator does not interact directly with the conversing parties.

Where are you getting that? Turing's most famous paper is just as Ukv describes. The link on that site doesn't work for me, but the reference is buried in their source: https://courses.cs.umbc.edu/471/papers/turing.pdf In Turing's test, the forced binary choice means P(human-judged-human) + P(machine-judged-human) is necessarily equal to 100%. This gives the 50% threshold clear intuitive and mathematical significance…

I don't understand why debates like this crop up. The premise of Turing's paper is plainly stated - that asking questions about an "Imitation Game" is more useful than asking whether machines can think.

That's all! He doesn't make any claim that the the game must be administered a particular way. In fact he spends only a few casual sentences glossing over how it would operate, and he's clearly just conveying the idea in broad strokes, not trying to describe an experimental procedure. And he says nothing at all about how the results might be judged, let alone thresholds for anything.

The paper is about what sorts of questions we should examine, not about specifically how they should be examined. So it seems weird to consider a test "bastardized" just because it doesn't match how you interpret Turing's casual description.

Re: Killed by LLM

#83
post #76

Earlier quoted context omitted.

Where are you getting that? Turing's most famous paper is just as Ukv describes. The link on that site doesn't work for me, but the reference is buried in their source: https://courses.cs.umbc.edu/471/papers/turing.pdf In Turing's test, the forced binary choice means P(human-judged-human) + P(machine-judged-human) is necessarily equal to 100%. This gives the 50% threshold clear intuitive and mathematical significance…

It’s interesting that even though you link to the original paper, you still repeat a very common incorrect summary of the task. The interrogator is not required to judge which of A or B is human, they are required to judge which is a woman on the implicit (though incorrect, in the case of interest) assumption that A and B are both human. While this amounts to more or less the same thing, it’s an interesting nuance th…

This gets brought up a lot, but it seems to me like a simple misreading.

Turing describes an initial game with a man (A) and a woman (B), where A's goal is to imitate B, and then asks: "what will happen if a machine takes the place of A?" I suppose it's possible that he meant the machine takes A's place by imitating a woman, but it's a lot more plausible that he meant the machine takes A's place by imitating B, i.e. a person.

Also there are several quotes later on that make no sense under your reading - check out the quotes including "imitation of the behaviour of a man" and "the part of B being taken by a man". Those quotes (maybe others, I didn't look) only make sense if the game is for the machine to imitate a person, not a woman.

Re: Killed by LLM

#84
post #83
post #76

Earlier quoted context omitted.

It’s interesting that even though you link to the original paper, you still repeat a very common incorrect summary of the task. The interrogator is not required to judge which of A or B is human, they are required to judge which is a woman on the implicit (though incorrect, in the case of interest) assumption that A and B are both human. While this amounts to more or less the same thing, it’s an interesting nuance th…

This gets brought up a lot, but it seems to me like a simple misreading. Turing describes an initial game with a man (A) and a woman (B), where A's goal is to imitate B, and then asks: "what will happen if a machine takes the place of A?" I suppose it's possible that he meant the machine takes A's place by imitating a woman, but it's a lot more plausible that he meant the machine takes A's place by imitating B , i.e.…

I agree with your last paragraph (see my last paragraph). But I think the most natural reading of the initial task description is that the machine also pretends to be a woman. The line “we do not wish to penalize a machine for being unable to shine in beauty competitions” supports this interpretation, given that a beauty competition is an event for women, under the assumptions of the time. So I think there are conflicting cues in the paper as to the intended interpretation.

As you say in your other comment, though, I don’t think Turing thought the exact details of the game were important - which explains why he didn’t trouble to spell them out very exactly.

If I had to guess, I’d say that Turing assumes that as the machine has no gender, the only relevant difference between the machine and the woman is that one is human and one is not. So for the rest of the paper he focuses on that difference and is vague on the gendered aspect of the task.

Re: Killed by LLM

#85
post #84
post #83

Earlier quoted context omitted.

This gets brought up a lot, but it seems to me like a simple misreading. Turing describes an initial game with a man (A) and a woman (B), where A's goal is to imitate B, and then asks: "what will happen if a machine takes the place of A?" I suppose it's possible that he meant the machine takes A's place by imitating a woman, but it's a lot more plausible that he meant the machine takes A's place by imitating B , i.e.…

I agree with your last paragraph (see my last paragraph). But I think the most natural reading of the initial task description is that the machine also pretends to be a woman. The line “we do not wish to penalize a machine for being unable to shine in beauty competitions” supports this interpretation, given that a beauty competition is an event for women, under the assumptions of the time. So I think there are confli…

Um. I follow you but that's a pretty huge stretch, considering that nothing in the paper is inconsistent with the conventional reading (that the machine is to imitate a person). There are sentences that are consistent with other readings, but none that's inconsistent with the usual one.

Re: Killed by LLM

#86
post #85
post #84

Earlier quoted context omitted.

I agree with your last paragraph (see my last paragraph). But I think the most natural reading of the initial task description is that the machine also pretends to be a woman. The line “we do not wish to penalize a machine for being unable to shine in beauty competitions” supports this interpretation, given that a beauty competition is an event for women, under the assumptions of the time. So I think there are confli…

Um. I follow you but that's a pretty huge stretch, considering that nothing in the paper is inconsistent with the conventional reading (that the machine is to imitate a person). There are sentences that are consistent with other readings, but none that's inconsistent with the usual one.

The machine is imitating a person on both understandings of the task. The difference lies in C’s task (whether C is trying to find which of A and B is a woman and which is a man, or trying to find which is human and which is a machine).

I think the initial description of the task is genuinely ambiguous. Your interpretation of it hadn’t occurred to me before, but I do see it now. I still think that “…when a machine takes the part of A in this game…” is most naturally interpreted as leaving the task unaltered but for the man being replaced by a machine, rather than implicitly describing the task mutadis mutandis. But reasonable people can certainly differ on such questions of interpretation.

Honestly I think Turing’s whole framing of the task is unnecessarily elaborate and confusing. Why even bother describing the man/woman task to begin with? I am not sure. Popular descriptions of the ‘Turing Test’ don’t seem to find this framing of any expositionary value.

Re: Killed by LLM

#87
post #78

Earlier quoted context omitted.

That's not the original Turing test either. The original imitation game as proposed by Turing involves reading a text transcript of a human and a computer and having the evaluator determine which is which. The evaluator does not interact directly with the conversing parties.

The original Turing game is whether machine can pretend to be a woman better than a man can (via teletype) as judged by an interrogator: > We now ask the question, "What will happen when a machine takes the part of A in this game?" Will the interrogator decide wrongly as often when the game is played like this as he does when the game is played between a man and a woman? These questions replace our original, "Can mac…

"Man pretending to be woman vs real woman" was just an example used to introduce the question in the form of a party game between humans before moving onto the actual question of a machine pretending to be human vs a real human.

At a stretch, by looking at only your quoted snippet you could read that the machine is pretending to be a woman - but that interpretation is not consistent with the rest of the paper. For instance, "the best strategy [for the machine] is to try to provide answers that would naturally be given by a man".

Re: Killed by LLM

#88
post #86
post #85

Earlier quoted context omitted.

Um. I follow you but that's a pretty huge stretch, considering that nothing in the paper is inconsistent with the conventional reading (that the machine is to imitate a person). There are sentences that are consistent with other readings, but none that's inconsistent with the usual one.

The machine is imitating a person on both understandings of the task. The difference lies in C’s task (whether C is trying to find which of A and B is a woman and which is a man, or trying to find which is human and which is a machine). I think the initial description of the task is genuinely ambiguous. Your interpretation of it hadn’t occurred to me before, but I do see it now. I still think that “…when a machine ta…

I think the point of the man/woman version of the game is that it lets Turing propose his question in relative terms. He doesn't ask "can the machine fool somebody N% of the time?" (as several in this thread imagine), but rather "can the machine fool somebody as often as one person fools another under similar conditions?".

> Your interpretation of it hadn’t occurred to me before,

The idea of a Turing Test is pretty widely understood to mean a test where a person guesses which responses come from a machine, not where they guess someone's gender. So my interpretation here is just that the paper says what most people think it says.

Re: Killed by LLM

#89
post #88
post #86

Earlier quoted context omitted.

The machine is imitating a person on both understandings of the task. The difference lies in C’s task (whether C is trying to find which of A and B is a woman and which is a man, or trying to find which is human and which is a machine). I think the initial description of the task is genuinely ambiguous. Your interpretation of it hadn’t occurred to me before, but I do see it now. I still think that “…when a machine ta…

I think the point of the man/woman version of the game is that it lets Turing propose his question in relative terms. He doesn't ask "can the machine fool somebody N% of the time?" (as several in this thread imagine), but rather "can the machine fool somebody as often as one person fools another under similar conditions?". > Your interpretation of it hadn’t occurred to me before, The idea of a Turing Test is pretty w…

Sorry, I am being ambiguous myself. I meant that your interpretation of this specific sentence had not occurred to me:

> We now ask the question, "What will happen when a machine takes the part of A in this game?"

I think you are right about what Turing meant. But it had honestly never occurred to me before that this description of the game could be understood as a description of the standard 'Turing test'. So, for this reason, I had always been sympathetic to the point that the standard Turing test does not appear to be the test that Turing describes in the original paper.

Here is a paper that makes your case, in case anyone finds it interesting. https://www.researchgate.net/profile/Gualtiero-Piccinini/pub...

Re: Killed by LLM

#90
post #82

Earlier quoted context omitted.

Where are you getting that? Turing's most famous paper is just as Ukv describes. The link on that site doesn't work for me, but the reference is buried in their source: https://courses.cs.umbc.edu/471/papers/turing.pdf In Turing's test, the forced binary choice means P(human-judged-human) + P(machine-judged-human) is necessarily equal to 100%. This gives the 50% threshold clear intuitive and mathematical significance…

I don't understand why debates like this crop up. The premise of Turing's paper is plainly stated - that asking questions about an "Imitation Game" is more useful than asking whether machines can think. That's all! He doesn't make any claim that the the game must be administered a particular way. In fact he spends only a few casual sentences glossing over how it would operate, and he's clearly just conveying the idea…

For Turing's test with the binary choice, the pass threshold is clear. If the machine and human are indistinguishable, then the probabilities that they're judged human must be equal. Since they sum to 100%, they must both equal 50%, making that a meaningful pass threshold. (A slightly higher pass threshold should be used in practice for statistical convenience, since infinitely many trials are required to make a confidence interval exactly include 50%. I'd guess that's why Turing mentions 70% in his paper.)

Without the binary choice, what do you think is the correct pass threshold? Those probabilities can now sum to anything. For GPT-4 in Jones and Bergen's paper they sum to 121%, though please nobody say 60.5%. The threshold now obviously depends on the interrogator's prior--I'd judge very differently if I were told the witnesses were 99% human than if I were told they were 1% human.

In that paper, do you think the interrogators knew their witness had only a 25% chance of being human? If so, why? If not, how do you think that affected the result? In aggregate over all the witnesses, their interrogators seem to have judged correctly only 60% of the time, while always guessing "machine" would have scored 75%. How did they manage to score worse than chance?

Turing's formulation is elegant, admitting meaningful statistical analysis with minimum assumptions. Most modifications are not, and that paper's is particularly bad. Turing's description may seem casual, but it's filled with mathematical depth that should not be missed.

Post reply on HN