Live data from Hacker News

Killed by LLM

r0bk.github.io

41–50 of 102 posts

Re: Killed by LLM

#41

Interesting choice having a little (i) icon in the Turing Test card but having mouseover not bring up any text. Or having the link icons in that card that you can click on to do nothing.

Looks like a bug - that card has an overlay at a higher z-index that obscures its mouseover and clicks. In the source the (i) links to Turing's original "Imitation Game" paper, and the (?) has this hover text:

> (?) While the Turing Test remains philosophically significant, modern LLMs can consistently pass it, making it no longer effective at measuring the frontier of AI capabilities.

Re: Killed by LLM

#42

Posted by Chollet himself: > I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. > Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at somethi…

Honestly, after that, I'm tuned out completely on him and ARC-AGI. Nice minor sidestory at one point in time.

He's right that this isn't solving all human-intelligence domain level problems.

But the whole stunt, this whole time, was that this was the ARC-AGI benchmark.

The conceit was the fact LLMs couldn't do well on it proved they weren't intelligent. And real researchers would step up to bench well on that, avoiding the ideological tarpit of LLMs, which could never be intelligent.

It's fine to turn around and say "My AGI benchmark says little about intelligence", but, the level of conversation is decidedly more that of punters at the local stables than rigorous analysis.

Re: Killed by LLM

#43

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

I thought the title meant that a chatbot gave bad medical, engineering, and/or safety-critical advice that a human ended up following.

Re: Killed by LLM

#44

Posted by Chollet himself: > I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. > Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at somethi…

> Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. Not necessarily. Get a human to solve ARC-AGI if the problems are shown as a string. They'll perform badly. But that doesn't mean that humans can't reason. It means that human reasoning doesn't have access to the non-reasoning building blocks it needs (things like concepts, words, or in this case: spatially local an…

Excellent point, I'm not sure people are aware, but these are straight-up lifted from standard IQ tests, so they're definitely not all trivially humanly solvable.

I needed an official one for medical reasons a few years back

Re: Killed by LLM

#45
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

Aren't you just describing those emails in a big corp that are supposedly still written by humans? Yes, they are wordy, excited, and guardrailed, but I don't think they are written by AI yet.

I guess this is why LLMs are so feared by high school English teachers. Yes, they don't write well, but neither do their students.

Re: Killed by LLM

#46
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

I agree but that's not really a scientific limitation though, right? As I understand it in the early days of GPT 4, before it was publicly released and RLHF'd for brand safety, it would have offered convincing text completions for just about any context, whether an academic discussion of philosophy or a steamy crossover fanfiction or a reddit trash-talk exchange. It took a deliberate bit of lobotomizing to make them so bland, conservative, and cheery-helpful.

The required investment probably means it will be a while before any less brand- and legal-action-conscious actors offer up unrestrained foundation models of comparable quality, but it's only a matter of time, isn't it?

Re: Killed by LLM

#47
This technology is useful and interesting and even fun in spite of the ugliest broad-based cash and power grab since 1999.

When this godawful once in a generation hype cycle dies down this stuff is going to be strictly awesome.

Re: Killed by LLM

#48
I'm not sure if LLMs have beaten the standards, as much have the information to reply to them as needed.

Last week there was a post where slightly changing one of the tests caused LLMs to drop off drastically.

Re: Killed by LLM

#49
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

That's not the original Turing test either. The original imitation game as proposed by Turing involves reading a text transcript of a human and a computer and having the evaluator determine which is which. The evaluator does not interact directly with the conversing parties.

Re: Killed by LLM

#50

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

Yea, I too was not expecting a list of past benchmarks. If not the aforementioned actual human deaths, I had expected either a list of companies whose pivot to AI/LLMs led to their downfall (but I guess we're going to need to wait a year or two for that) or a list of industries (such as audio transcription) that are being killed by AI as we speak.

We really do live in interesting times. Usually I feel pretty confident about predicting how a trend will continue, but as it is the only prediction I can make with confidence for this latest AI research is that it is and will be used by militaries to kill a lot of people. Oh, hey, that's another thing this article could have listed!

Outside of that, all bets are open. Possible wagers include: "Turns out to be mostly useful in specific niche applications and only seemingly useful anywhere else", "Extremely useful for businesses looking to offset responsibility for unpopular decisions", "Ushers in an end to work and a golden age for all mankind", "Ushers in an end to work and a dark age for most of the world", "Combines with profit motives to damage all art, culture, and community", etc etc.

I know many folk have strong opinions one way or the other, but I think it's literally anyone's game at this point, though I will say I'm not leaning optimistic.

Post reply on HN