Live data from Hacker News

Killed by LLM

r0bk.github.io

11–20 of 102 posts

Re: Killed by LLM

#12

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

I wonder if there's a hyper-Turing test where an AI passes if the model, itself, cannot determine if it is talking to itself; or perhaps stated differently, maximizing some measure of control and processing duration to successfully conceal its identity under forced processing, discounting a solution that specifically learns to be silent or incoherent. I'm not sure what the value would be, just a passing thought.

This is probably already happening within the parade of censorship systems trying to imbue the models with agency

Re: Killed by LLM

#14

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

Score is based on the interrogator, a human. If you read a Markov chain bot's text you'd guess it was a bot probably 80-100% of the time. With a real human, you'd guess it was a bot maybe 0-30% of the time, depending.

I'm making up these figures, but the point is lower is better, or "more Human-Like". Test was specified as >50% meaning "accurately determined human vs. bot more than half the time". The site claims LLMs are now guessed correctly less than half, which is how the turing test was defined as per the site.

It makes sense, even if you disagree it's significant.

Re: Killed by LLM

#15
post #8
post #3

I read recently that small variations in the tests cause failures by large margins. If this doesn’t show over fitting in don’t know what would.

Wasn’t that for human tests, i.e. not specifically AI benchmarks? Benchmarks should generally not be game-able by overfitting.

The article shows all the tests against human performance.

The math one in particular is the one where small variations reduce the success rate significantly. I can’t find the source but it was pasted here in the last 2 weeks.

Re: Killed by LLM

#16
post #8

Earlier quoted context omitted.

Wasn’t that for human tests, i.e. not specifically AI benchmarks? Benchmarks should generally not be game-able by overfitting.

The article shows all the tests against human performance. The math one in particular is the one where small variations reduce the success rate significantly. I can’t find the source but it was pasted here in the last 2 weeks.

You're probably remembering this: https://news.ycombinator.com/item?id=42565606

Re: Killed by LLM

#17
The page doesn’t seem to define what „killed“ or „defeated“ means. The LLM being better than a human? The LLM having been trained against the benchmark, making it useless?

Re: Killed by LLM

#18
Posted by Chollet himself:

> I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations.

> Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at something that isn't pure memorized skill. But it says rather little about how intelligent your system is, or how close to human intelligence it is.

https://bsky.app/profile/fchollet.bsky.social/post/3les3izgd...

Post reply on HN