Live data from Hacker News

A Man Out to Prove How Dumb AI Still Is

theatlantic.com

21–30 of 66 posts

Re: A Man Out to Prove How Dumb AI Still Is

#21

Earlier quoted context omitted.

I think it's fascinating that his impossible benchmark got defeated, but because the Keras guy doesn't like LLMs, it is possible to mishear algorithmic distaste as saying people shipping this are "lazy" and "hype maxing."

Francois never said he dislike LLMs. In fact, he said he expected them to be part of the solution to ARC. I don’t know where this persistent myth comes from, but it has to go.

> part of the solution to ARC. I don’t know where this persistent myth comes from,

Part of, explicitly, not the, quite 100% explicitly. The TL;DR is "LLMs can't do it alone, program synthesis leveraging LLMs is my bet". Not "Maybe not LLMs but they'll certainly help us get there!", quite the opposite! Hence: well, TFA. And the intellectually lazy quote we are explicitly discussing. And anything Chollet has said on the subject. [^1]

[^1]"LLMs won’t lead to AGI - $1,000,000 Prize to find true solution" - https://www.dwarkesh.com/p/francois-chollet - 1.5 hours with the gent

Re: A Man Out to Prove How Dumb AI Still Is

#22
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

> best solved with what used to be called symbolic AI before it started working

Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact.

Btw Chollet has said basically as much. He calls them “stored programs” I think.

I think he is onto something. The right atomic to approach these problems is probably not the token, at least at first. Higher level abstraction should be refined to specific components, similar to the concept of diffusion.

Re: A Man Out to Prove How Dumb AI Still Is

#23
> A person who scores 30 percent on ARC-AGI-2 is in no sense inferior to someone who scores 90 percent

"News just in: journalist for the Atlantic stops reasoning and drifts in a world of feelings after neural hijacking, as he perceives abilities as some kind of threat".

> Human cognitive diversity [...] when that diversity is already so abundant, do you really want to?

We definitely need intelligence.

Re: A Man Out to Prove How Dumb AI Still Is

#24
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

>> If you want to saturate any model today give it a string and a grammar and ask it to generate the string from the grammar. I'm not sure I understand what that means - could you explain please?

It means applying specific rules about how text can be generated. For example, generating valid json reliably. Currently we use constrained decoding to accomplish this (e.g. the next token must be one of three valid options).

Now you can imagine giving an LLM arbitrary validity rules for generating text. I think that’s what they mean by “grammar”.

Re: A Man Out to Prove How Dumb AI Still Is

#25
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.

Re: A Man Out to Prove How Dumb AI Still Is

#26
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

> that’s good enough as far as I’m concerned

But in that case, why an LLM. If we want Question-Answer machines to be reliable, they must have the skills which include "counting" just as a basic example.

Re: A Man Out to Prove How Dumb AI Still Is

#27
post #25

Earlier quoted context omitted.

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.

No LLM or other modern AI architecture I'm aware of is supposed to model the human brain. Even if they were, LLMs can add large numbers with the level of skill I'd expect from a 10 year old:

----

What's 494547645908151+7640745309351279642?

ChatGPT said: The sum of 494,547,645,908,151 and 7,640,745,309,351,279,642 is:

7,641,239,857,997,187,793

----

(7,641,239,856,997,187,793 is the correct answer)

Re: A Man Out to Prove How Dumb AI Still Is

#28

Not to dismiss Chollet’s work, but I’m starting to think he need prove nothing to even the muggles. For example, nearly any endurance athlete stands a good chance of being a Strava user. If you run in those circles, have you heard a single person with anything good to say about Strava’s “Athletic Intelligence”? Garmin is rolling out a beta right now that includes “AI Insights” or summat. Same deal: useless summaries…

[deleted]

Re: A Man Out to Prove How Dumb AI Still Is

#29
post #25

Earlier quoted context omitted.

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.

If I were to guess, most (adult) humans could not add two 3 digit numbers together with 100% accuracy. Maybe 99%? Computers can already do 100%, so we should probably be trying to figure out how to use language to extract the numbers from stuff and send them off to computers to do the calculations. Especially because in the real world most numbers that matter are not just two digits addition

Re: A Man Out to Prove How Dumb AI Still Is

#30

Earlier quoted context omitted.

I think it's fascinating that his impossible benchmark got defeated, but because the Keras guy doesn't like LLMs, it is possible to mishear algorithmic distaste as saying people shipping this are "lazy" and "hype maxing."

Arc agi 1 that "got defeated" was published even before first mainstream llms and still stood the test of time

> "got defeated"

1.5 hours with Chollet on "LLMs won’t lead to AGI - $1,000,000 Prize to find true solution" -

Published June 2024, and by December, well...we can all agree there's an ARC AGI 2 now.

https://www.dwarkesh.com/p/francois-chollet

Post reply on HN