Live data from Hacker News

A Man Out to Prove How Dumb AI Still Is

theatlantic.com

41–50 of 66 posts

Re: A Man Out to Prove How Dumb AI Still Is

#41
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.

Just copied your prompt and it handled it just fine.

Re: A Man Out to Prove How Dumb AI Still Is

#42
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.

With o3-mini-high (just the last paragraph):

civilization Mycenaean the of practices religious and economic, administrative the into insights invaluable provides and B Linear as known script the in recorded was language Greek the of form attested earliest The

Re: A Man Out to Prove How Dumb AI Still Is

#43
post #9

>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…

I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.

As much I think AI is overhyped too, that is a prime use case that would be better solved by passing the text to a tool, rather than jam a complex transformations like that into its latent space.

Re: A Man Out to Prove How Dumb AI Still Is

#44

Earlier quoted context omitted.

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

As soon as the companies behind these systems stop marketing them as do-anything machines, I will stop judging them on their ability to do everything. The ChatGPT input field still says ‘Ask anything’, and that is what I shall do.

You can ask me anything. I don’t see that as a promise that I am infallible.

Re: A Man Out to Prove How Dumb AI Still Is

#45
post #25

Earlier quoted context omitted.

> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…

Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.

Most human ten year olds can add two large numbers together with the aid of a scratchpad and a pen. You need tools other than a single dimensional vector of text to do some of these things.

Re: A Man Out to Prove How Dumb AI Still Is

#46
post #25

Earlier quoted context omitted.

Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.

No LLM or other modern AI architecture I'm aware of is supposed to model the human brain. Even if they were, LLMs can add large numbers with the level of skill I'd expect from a 10 year old: ---- What's 494547645908151+7640745309351279642? ChatGPT said: The sum of 494,547,645,908,151 and 7,640,745,309,351,279,642 is: 7,641,239,857,997,187,793 ---- (7,641,239,856,997,187,793 is the correct answer)

I tried it on gpt-4-turbo and it seems to give the right answer:

>Let's calculate:494,547,645,908,151+7,640,745,309,351,279,642=7,641,239,856,997,187,793 >494,547,645,908,151+7,640,745,309,351,279,642=7,641,239,856,997,187,793 >Answer: 7,641,239,856,997,187,793

Re: A Man Out to Prove How Dumb AI Still Is

#47

Not to dismiss Chollet’s work, but I’m starting to think he need prove nothing to even the muggles. For example, nearly any endurance athlete stands a good chance of being a Strava user. If you run in those circles, have you heard a single person with anything good to say about Strava’s “Athletic Intelligence”? Garmin is rolling out a beta right now that includes “AI Insights” or summat. Same deal: useless summaries…

I think some of the AI demos are kind of comedy gems.

I have seen the Apple Intelligence presentation a while ago and in the span of five minutes they had someone asking the assistant to expand a one liner into an e-mail and then someone receiving a long e-mail and asking the AI to summarize it.

We spun GPUs to expand, then spun them again to summarize. Gold.

Re: A Man Out to Prove How Dumb AI Still Is

#48
post #41

Earlier quoted context omitted.

I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.

Just copied your prompt and it handled it just fine.

?siht ekil kool rewsna eht diD

Edit: realized just now that my summary of the 'test' failed to specify the request fully: the letters need to be reversed, too. Maybe I'm just bad with AI tools, because I didn't even get a response that 'this like looked' (i.e. reversed the order of the words).

Re: A Man Out to Prove How Dumb AI Still Is

#49
post #42

Earlier quoted context omitted.

I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.

With o3-mini-high (just the last paragraph): civilization Mycenaean the of practices religious and economic, administrative the into insights invaluable provides and B Linear as known script the in recorded was language Greek the of form attested earliest The

Oh, interesting, what do you get when you specify that the letters need to be reversed, too? (That was what I meant and the original prompt explicitly stated that requirement. I forgot to include it in the summary of my 'test' here.)

Re: A Man Out to Prove How Dumb AI Still Is

#50

Not to dismiss Chollet’s work, but I’m starting to think he need prove nothing to even the muggles. For example, nearly any endurance athlete stands a good chance of being a Strava user. If you run in those circles, have you heard a single person with anything good to say about Strava’s “Athletic Intelligence”? Garmin is rolling out a beta right now that includes “AI Insights” or summat. Same deal: useless summaries…

I'm taking a university course right now and one of the big textbook companies (McGraw-Hill, Macmillan, one of those) has an "AI Tutor" on their homework assignment.

It's somehow LESS helpful than just having a pointer to which part of the text to revisit.

It basically just restates the question/problem. Even worse than that is that it's an essentially STATIC note for each question yet it appears to be REAL-TIME GENERATED each time. I guess that could just be for appearances but it's just dumb all the way around really.

Post reply on HN