Live data from Hacker News

Killed by LLM

r0bk.github.io

1–10 of 102 posts

Re: Killed by LLM

#3
I read recently that small variations in the tests cause failures by large margins.

If this doesn’t show over fitting in don’t know what would.

Re: Killed by LLM

#4
How does this site make sense?

It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%.

At that point I just stopped scrolling.

Re: Killed by LLM

#5

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

the Turing test is scored by how often an interrogator can determine if they are talking to a machine or a human. It’s perhaps a confusing way to show it, and leaves out a lot of important information about the result they are citing, but they are saying before the interrogator did better than chance and after gpt the interrogator guesses right slightly worse than chance.

Re: Killed by LLM

#6

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

Presumably it means that the human detected the AI correctly less than 50% of the time, averaged over a repeated number of experiments.

Re: Killed by LLM

#7

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

I read that as the pass mark being able to identify human vs machine more than 50% of the time. At 50% it's no better than randomly guessing.

Re: Killed by LLM

#8
post #3

I read recently that small variations in the tests cause failures by large margins. If this doesn’t show over fitting in don’t know what would.

Wasn’t that for human tests, i.e. not specifically AI benchmarks? Benchmarks should generally not be game-able by overfitting.

Re: Killed by LLM

#9

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

[deleted]

Re: Killed by LLM

#10
The layer with the radial gradient you're putting in front of the Turing Test card blocks interaction with it - can't click or hover on its links.
Post reply on HN