Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
1–10 of 62 posts
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#2The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model to spit out the jargon I need.
However, the moment I ask a real question that goes beyond summarizing something which is covered thousands of times online, I am immediately let down.
Is this just a result of the foundation of the model being the world best autocompletion engine? My assessment is “yes” and I don’t believe that any of the modifications coming, like plugins, will fundamentally change this.
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#3> ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4
But reading the paper the challenges it is failing on are ones that I wager the average human would fail on too (at least a good portion of the time).
The paper might strictly be accurate, but I think we should try and bring these papers back to a real-world context - which is that it’s probably operating above your average human at these tasks.
Is superhuman/genius-level capability really required before we say the LLMs are any good?
(I see this view on HN too - statements like ‘LLMs can’t create novel maths theorems!’ as an argument that LLMs aren’t good at reasoning, disregarding that most humans today can’t find novel/undiscovered maths theorems)
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#4I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#5I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
People want it to perform better than any expert human at any possible subject before it's considered "real AI". It isn't enough for critics for it to be better than the average person at virtually everything its put to the test on.
It seems like there is some resentment and almost anger at this technology, particularly with the artistic AIs like Midjourney. I can understand that more readily, but what's the real beef with ChatGPT?
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#6I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#7This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#8This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…
Furthermore, I just don’t feel like the transformer architecture is suited for problem solving. Like I may just be a charlatan but self attention over the space of words does not seem like it’s going to be enough, and praying it falls out in emergent behavior if we can just add more parameters is… unscientific-ish? Now, if you could figure out a way to do self-attention over the space of concepts? Maybe you’ve got something.
I feel like AlphaGo ideas and some variation on MCTS is more likely to produce a solid problem solving architecture.
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#9I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
The goal posts for AI are moving quickly, and in my mind, a lot of the criticism os too shallow. People want it to perform better than any expert human at any possible subject before it's considered "real AI". It isn't enough for critics for it to be better than the average person at virtually everything its put to the test on. It seems like there is some resentment and almost anger at this technology, particularly w…
And many of those things are worse at their task than a person except for speed and scaling. Can a machine be fooled by dazzle for recognizing a face in a way a human can't? Sure, but nobody is willing to pay for a human to go through everyone's photo albums...
But does ChatGPT "use AI" as a tool in the same sense that Spotify's recommendations "use AI" or is it "an AI" in the sense that it's an independent consciousness/agent?
This is the first time so many people have disagreed on that part. And that skews the debate into "a person is better" vs talking about if a person is even practical in most of the situations we'd use this.
Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
#10I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
The goal posts for AI are moving quickly, and in my mind, a lot of the criticism os too shallow. People want it to perform better than any expert human at any possible subject before it's considered "real AI". It isn't enough for critics for it to be better than the average person at virtually everything its put to the test on. It seems like there is some resentment and almost anger at this technology, particularly w…
We don't exactly know what it can and can't do, a property which in a computer program at any rate is deeply mysterious and unusual. It initially gives the appearance of being a human which knows everything. This leads a lot of people to angrily declare that its appearance is deceptive, and in searching for words to describe in exactly what way it falls short, they incorporate flawed intuition on what it is capable of. So there's a lot of back and forth right now as we collectively swap memes to try and make sense of such a dramatic development.