>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
A Man Out to Prove How Dumb AI Still Is
41–50 of 66 posts
Re: A Man Out to Prove How Dumb AI Still Is
#42>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
civilization Mycenaean the of practices religious and economic, administrative the into insights invaluable provides and B Linear as known script the in recorded was language Greek the of form attested earliest The
Re: A Man Out to Prove How Dumb AI Still Is
#43>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
Re: A Man Out to Prove How Dumb AI Still Is
#44Earlier quoted context omitted.
> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…
As soon as the companies behind these systems stop marketing them as do-anything machines, I will stop judging them on their ability to do everything. The ChatGPT input field still says ‘Ask anything’, and that is what I shall do.
Re: A Man Out to Prove How Dumb AI Still Is
#45Earlier quoted context omitted.
> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…
Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.
Re: A Man Out to Prove How Dumb AI Still Is
#46Earlier quoted context omitted.
Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.
No LLM or other modern AI architecture I'm aware of is supposed to model the human brain. Even if they were, LLMs can add large numbers with the level of skill I'd expect from a 10 year old: ---- What's 494547645908151+7640745309351279642? ChatGPT said: The sum of 494,547,645,908,151 and 7,640,745,309,351,279,642 is: 7,641,239,857,997,187,793 ---- (7,641,239,856,997,187,793 is the correct answer)
>Let's calculate:494,547,645,908,151+7,640,745,309,351,279,642=7,641,239,856,997,187,793 >494,547,645,908,151+7,640,745,309,351,279,642=7,641,239,856,997,187,793 >Answer: 7,641,239,856,997,187,793
Re: A Man Out to Prove How Dumb AI Still Is
#47Not to dismiss Chollet’s work, but I’m starting to think he need prove nothing to even the muggles. For example, nearly any endurance athlete stands a good chance of being a Strava user. If you run in those circles, have you heard a single person with anything good to say about Strava’s “Athletic Intelligence”? Garmin is rolling out a beta right now that includes “AI Insights” or summat. Same deal: useless summaries…
I have seen the Apple Intelligence presentation a while ago and in the span of five minutes they had someone asking the assistant to expand a one liner into an e-mail and then someone receiving a long e-mail and asking the AI to summarize it.
We spun GPUs to expand, then spun them again to summarize. Gold.
Re: A Man Out to Prove How Dumb AI Still Is
#48Earlier quoted context omitted.
I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
Just copied your prompt and it handled it just fine.
Edit: realized just now that my summary of the 'test' failed to specify the request fully: the letters need to be reversed, too. Maybe I'm just bad with AI tools, because I didn't even get a response that 'this like looked' (i.e. reversed the order of the words).
Re: A Man Out to Prove How Dumb AI Still Is
#49Earlier quoted context omitted.
I tried my own test recently: "Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top." ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
With o3-mini-high (just the last paragraph): civilization Mycenaean the of practices religious and economic, administrative the into insights invaluable provides and B Linear as known script the in recorded was language Greek the of form attested earliest The
Re: A Man Out to Prove How Dumb AI Still Is
#50Not to dismiss Chollet’s work, but I’m starting to think he need prove nothing to even the muggles. For example, nearly any endurance athlete stands a good chance of being a Strava user. If you run in those circles, have you heard a single person with anything good to say about Strava’s “Athletic Intelligence”? Garmin is rolling out a beta right now that includes “AI Insights” or summat. Same deal: useless summaries…
It's somehow LESS helpful than just having a pointer to which part of the text to revisit.
It basically just restates the question/problem. Even worse than that is that it's an essentially STATIC note for each question yet it appears to be REAL-TIME GENERATED each time. I guess that could just be for appearances but it's just dumb all the way around really.