Live data from Hacker News

Killed by LLM

r0bk.github.io

21–30 of 102 posts

Re: Killed by LLM

#21
Interesting choice having a little (i) icon in the Turing Test card but having mouseover not bring up any text. Or having the link icons in that card that you can click on to do nothing.

Re: Killed by LLM

#23
post #3

I read recently that small variations in the tests cause failures by large margins. If this doesn’t show over fitting in don’t know what would.

Eventually, all the better AGI tests should have large private evaluation datasets with no possible cheating or feedback loops. We're getting there.

Re: Killed by LLM

#25

Posted by Chollet himself: > I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. > Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at somethi…

> Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations.

Not necessarily. Get a human to solve ARC-AGI if the problems are shown as a string. They'll perform badly. But that doesn't mean that humans can't reason. It means that human reasoning doesn't have access to the non-reasoning building blocks it needs (things like concepts, words, or in this case: spatially local and useful visual representations).

Humans have good resolution-invariant visual perception. For example, take an ARC-AGI problem, and for each square, duplicate it a few times, increasing its resolution from X*X to 2X*2X. To a human, the problem will be almost exactly equally difficulty. Not for LLMs that have to deal with 4x as much context. Maybe for an LLM if it can somehow reason over the output of a CNN, and if it was trained to do that like how humans are built to do that.

Re: Killed by LLM

#26

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

From TFA:

>GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%).

Re: Killed by LLM

#27
I don't really understand why "Killed by: Saturation" is needed - what other options are there?

It would also be nice to see the "unbeaten" list: standardized tests LLMs still fail (for now). e.g. Wozniak's coffee test.

Re: Killed by LLM

#28

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

Similarly I thought it would be about ML and data projects that have become defunct due to the advent of LLMs.

Re: Killed by LLM

#29
The tortoise lays on its back, its belly baking in the hot sun, beating its legs trying to turn itself over, but it can't. Not without your help. But you're not helping.

Re: Killed by LLM

#30

The tortoise lays on its back, its belly baking in the hot sun, beating its legs trying to turn itself over, but it can't. Not without your help. But you're not helping.

Describe in single words, only the good things that come into your mind about your mother.
Post reply on HN