I tried the prime number problem and GPT-4 nailed it. I’m not sure whether they are testing things correctly… “ Sure, let's go step by step. A prime number is a number greater than 1 that has no positive divisors other than 1 and itself. This means if we can find any other number (excluding 1 and the number itself) that divides 17077, then it is not a prime number. Let's start by checking divisibility by 2. Since 170…
Can you define "nailing" it? Cause this isn't what I'd call nailing it...
How is ChatGPT's behavior changing over time?
151–160 of 187 posts
Re: How is ChatGPT's behavior changing over time?
#152Earlier quoted context omitted.
In my experience, it can reason usably well but acts like it has dyscalculia. It'll set up a proper algorithm, step through it and trip over digits.
"It'll set up a proper algorithm, step through it and trip over digits." Or it is just pretending to do so. And since it pretends, of course it trips over all small things as it does not understand them.
And in this case, the shape of the answer is often right; it just makes ... ordinary errors. Ironically, the AI is a lot better at high-level thinking than correct calculation.
Re: How is ChatGPT's behavior changing over time?
#153Earlier quoted context omitted.
That is very much not random, but generally enough for 99% of all use cases.
I would not bet on whether chatgpt's results would be more or less random than this process.
Re: How is ChatGPT's behavior changing over time?
#154This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' task selection and evaluation method and can easily result from fine tuning rather than intentionally degrading GPT-4's performance for cost saving reasons. They report 2 degradations: code generation & math problems. In both cases, they report a behavior change (likely fine tuning) rather than a capability decrease (p…
In my opinion the more likely thing is that OpenAI is gaslighting people that the finetuning is improving the model when it likely mostly improves safety at some cost to capability. I'd bet this is measured against a set of evals and it looks like it performs well BUT I'd also bet the evals are asymmetrically good at detecting "unsafe" or jailbreak behavior and bad at detecting reduced general cognitive flexibility.…
This is not necessarily the case, and even if it is doesn't imply gaslighting as compared to inability to measure.
Re: How is ChatGPT's behavior changing over time?
#155Earlier quoted context omitted.
"We" are never going to stop trying to use LLMs for math. They are obviously mimicking a smart person. What do you do with smart persons? You query them with bad and lazy questions (often without being too honest about that to yourself), and hope/expect helpful answers. Because that is how it works. The limiting factor so far has been the smart persons time and patience. Now, no longer. People moderating their LLMs u…
>"We" are never going to stop trying to use LLMs for math. They are obviously mimicking a smart person. A lot of words to make a big deal out of nothing. All that is needed is some new abstracted layer that identifies a math question and then proxies it over to the wolfram plugin. That’s it We don’t have crazy debates over whether a polygon should be rendered by the cpu or a gpu. We solved this problem
Re: How is ChatGPT's behavior changing over time?
#156Earlier quoted context omitted.
We should also test its capabilities on cooking steak, flying rockets, and making love. Only then will we know if AI can be superior to humans on all things.
Well, if there was an API hooked up to a webcam & 6-dof arm that would be an interesting task. (The steak)
Re: How is ChatGPT's behavior changing over time?
#157Earlier quoted context omitted.
What motive do you see for them releasing these models for free, especially with the massive increase in cost associated with catching up?
Kill Op*nAI, for starter. If they see it as a threat, commoditizing the tech is a great way to get rid of them at a reasonable cost.
Re: How is ChatGPT's behavior changing over time?
#158Earlier quoted context omitted.
It’s not about the results, it’s about its ability to “reason”. Math is about as close to pure reasoning we get so I don’t get the pessimism. If it is bad at math and can’t be taught, then you have a fundamental problem. It’s a matter of time before this limit gets hit in other domains.
In my experience, it can reason usably well but acts like it has dyscalculia. It'll set up a proper algorithm, step through it and trip over digits.
100,000 + 987 - 1444 * 25,945.842 / 0.0042
becomes
"100" one hundred
"," comma
"000" triple zero
" +" space plus
" 9" space nine
"87" eighty seven
" -" space minus
" 14" space 14
"44" forty four
" *" space times
" 25" space twenty five
"," comma
"9" nine
"45" forty five
"." period
"8" eight
"42" forty two
" /" space divide
" 0" zero
"." period
"00" double zero
"42" forty two
Now imagine someone reading that to you over the phone once and asking you to do the math in your head and you aren't allowed to use paper and pencil and you have to get it right the first time.
Re: How is ChatGPT's behavior changing over time?
#159I think we should stop trying to quiz LLMs on mathematics, something for which they are explicitly not designed to do with their tokenized view of the world. Ask GPT-4 to use its Wolfram plugin and it returns the answers quickly and correctly. Second, I think the code generation bit of this paper is blown out of proportion. The code can't be immediately injected into a codebase due to a formatting change (triple quot…
Seriously. GPT doing math is like using a 737 to drive around on the ground, or if you had the phone number of a prominent astrophysicist and you call him to do long division for you. Wtf is the point. We have computer things to do every math problem. It’s a waste of energy to use LLMs for it in my opinion.
Re: How is ChatGPT's behavior changing over time?
#160Earlier quoted context omitted.
> Seriously. GPT doing math is like using a 737 to drive around on the ground, or if you had the phone number of a prominent astrophysicist and you call him to do long division for you. Wtf is the point. We have computer things to do every math problem. It’s a waste of energy to use LLMs for it in my opinion. Using GPT to do maths is probably like using a 737 to drive around on the ground. Teaching GPT to do maths mi…
Well that's... a claim. What's your reasoning for operations on digits within the model itself improving any current metrics?
Humans are different to an AI, but putting that aside, my intuition would be that if we never taught kids any mental maths, their concept/understanding of numbers would be fundamentally different to how it is if they learn that 9 x 9 = 81 (also look at how your fingers move - there is a relationship there!).
But who knows, AI is strange and there's lots of stuff that needs to be experimented with. I would think training an intuitive sense of numbers would have other fall-outs though. This is half the beauty of LLM's right? You show an LLM some history books and it also learns about biology, politics, grammar, etymology and love. You teach a LLM maths and it also learns ... ?