Earlier quoted context omitted.
> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…
> that’s good enough as far as I’m concerned But in that case, why an LLM. If we want Question-Answer machines to be reliable, they must have the skills which include "counting" just as a basic example.
A Man Out to Prove How Dumb AI Still Is
31–40 of 66 posts
Re: A Man Out to Prove How Dumb AI Still Is
#32Earlier quoted context omitted.
> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…
Most human ten year olds in school can add two large numbers together. If a connectionist network is supposed to model the human brain, it should be able to do that. Maybe LLMs can do a lot of things, but if they can't do that, then they're an incomplete model of the human brain.
For what it’s worth, people are also pretty bad at math compared to calculators. We are slow and error prone. That’s ok.
What I was (poorly) trying to say is that I don’t care if the neural net solves the problem if it can outsource it to a calculator. People do the same thing. What is important is reliably accomplishing the goal.
Re: A Man Out to Prove How Dumb AI Still Is
#33>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
Sounds like a use-case for property testing: https://en.wikipedia.org/wiki/Software_testing#Property_test...
Re: A Man Out to Prove How Dumb AI Still Is
#34> To hit 87 percent on the original ARC-AGI test, o3 spent roughly 14 minutes per puzzle and, by my calculations, may have required hundreds of thousands of dollars in computing and electricity > the bot came up with more than 1,000 possible answers per grid before selecting a final submission. Yeah, AGI is right around the corner… /s
Re: A Man Out to Prove How Dumb AI Still Is
#35Re: A Man Out to Prove How Dumb AI Still Is
#36>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
> best solved with what used to be called symbolic AI before it started working Right, the current paradigm of requiring an LLM to do arbitrary digit multiplication will not work and we shouldn’t need to. If your task is “do X” and it can be reliably accomplished with “write a python program to do X” that’s good enough as far as I’m concerned. It’s preferable, in fact. Btw Chollet has said basically as much. He calls…
The ChatGPT input field still says ‘Ask anything’, and that is what I shall do.
Re: A Man Out to Prove How Dumb AI Still Is
#371a) it's not AI, it's LLM. The companies who create/train/operate them may (wink-wink) pitch them as "AI" with half-truths, but we (here) know it's LLMs "all the way down" 1b) just like I disliked the "autopilot" in Teslas because it was never autopilot. 2) I know that I wanted to write some software tools, and I have been successful at this for the past many months, and I got top-shelve tools, that work, do their ta…
I completely agree, they are a tool, and a decently useful tool. They are not early in the curve, they’re about flat at this point.
Re: A Man Out to Prove How Dumb AI Still Is
#38>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
Show me the results of your symbolic AI on ARC 2.
Re: A Man Out to Prove How Dumb AI Still Is
#39>Last week, the ARC Prize team released an updated test, called ARC-AGI-2, and it appears to have sent the AIs back to the drawing board. The full o3 model has not yet been tested, but a version of o1 dropped from 32 percent on the original puzzles to just 3 percent on the new version, and a “mini” version of o3 currently available to the public dropped from roughly 30 percent to below 2 percent. (An OpenAI spokesper…
"Write a history of the Greek language but reverse it, so that one would need to read it from right to left and bottom to top."
ChatGPT wrote the history and showed absolutely no awareness, let alone, "understanding" of the second half of the prompt.
Re: A Man Out to Prove How Dumb AI Still Is
#40https://chatgpt.com/share/67ef43f4-3b88-800d-a5a3-e3ffea178f...
(Me trying to describe a desk top with a fold down hinged top, and it just drawing whatever)