Killed by LLM
21–30 of 102 posts
Re: Killed by LLM
#22The page doesn’t seem to define what „killed“ or „defeated“ means. The LLM being better than a human? The LLM having been trained against the benchmark, making it useless?
Re: Killed by LLM
#23I read recently that small variations in the tests cause failures by large margins. If this doesn’t show over fitting in don’t know what would.
Re: Killed by LLM
#24Everything says "killed by saturation". Is there another way to be killed?
Re: Killed by LLM
#25Posted by Chollet himself: > I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. > Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at somethi…
Not necessarily. Get a human to solve ARC-AGI if the problems are shown as a string. They'll perform badly. But that doesn't mean that humans can't reason. It means that human reasoning doesn't have access to the non-reasoning building blocks it needs (things like concepts, words, or in this case: spatially local and useful visual representations).
Humans have good resolution-invariant visual perception. For example, take an ARC-AGI problem, and for each square, duplicate it a few times, increasing its resolution from X*X to 2X*2X. To a human, the problem will be almost exactly equally difficulty. Not for LLMs that have to deal with 4x as much context. Maybe for an LLM if it can somehow reason over the output of a CNN, and if it was trained to do that like how humans are built to do that.
Re: Killed by LLM
#26How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.
>GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%).
Re: Killed by LLM
#27It would also be nice to see the "unbeaten" list: standardized tests LLMs still fail (for now). e.g. Wozniak's coffee test.
Re: Killed by LLM
#28I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...
Re: Killed by LLM
#29Re: Killed by LLM
#30The tortoise lays on its back, its belly baking in the hot sun, beating its legs trying to turn itself over, but it can't. Not without your help. But you're not helping.