Live data from Hacker News

ChatGPT-4o vs. Math

sabrina.dev

141–150 of 182 posts

Re: ChatGPT-4o vs. Math

#141
post #96
post #21

I posted the same 'Zero-Shot Chain-of-Thought and Image' to ChatGPT-4o and it made the same error. I then followed up with 'Your math is good but you derived incorrect data from the image. Can you take another look and see if you can tell where the error is?'. It figured it out and corrected it: Let's re-examine the image and the data provided: * The inner radius r1 is given as 5cm * The outer radius r2 is given as 1…

Once you correct the LLM, it will continue to provide the corrected answer until some time later, when it will again make the same mistake. At least, this has been my experience. If you are using LLM to pull answers programmatically and rely on their accuracy, here is what worked for the structured or numeric answers, such as numbers, JSON, etc. 1) Send the same prompt twice, including "Can you double check?" in the…

via api (harder to do via chat as cleanly) you can also try showing it do a false attempt (but a short one so it's effectively part of the prompt) and then you say try again.

Re: ChatGPT-4o vs. Math

#142
post #139

Earlier quoted context omitted.

I like your theory but if it’s true, then Ilya was wrong. All of the current LLM architectures have no medium-term memory or iterative capability. That means they’re missing essential functionality for general intelligence. I tired GPT 4o for various tasks and it’s good but it isn’t blowing my skirt up. The only noticeable difference is the speed, which is a very nice improvement that enables new workflows.

Part of the confusion is that people use the term "AGI" to mean different things. We should actually call this AGI, because it is starkly different from the narrow capabilities of AI a few years ago. I am not claiming that it is a full digital simulation of a human being or has all of the capabilities of animals like humans, or is the end of intelligence research. But it is obviously very general purpose at this poin…

Currently, they’re like Dory from Finding Nemo: long and short term memory but they forget everything after each conversation.

The character of Dory is jarring and bizarre precisely because of this trait! Her mind is obviously broken in a disturbing way. AIs give me the same feeling. Like talking to an animatronic robot at a theme park or an NPC in a computer game.

Re: ChatGPT-4o vs. Math

#143
post #139

Earlier quoted context omitted.

Part of the confusion is that people use the term "AGI" to mean different things. We should actually call this AGI, because it is starkly different from the narrow capabilities of AI a few years ago. I am not claiming that it is a full digital simulation of a human being or has all of the capabilities of animals like humans, or is the end of intelligence research. But it is obviously very general purpose at this poin…

Currently, they’re like Dory from Finding Nemo: long and short term memory but they forget everything after each conversation. The character of Dory is jarring and bizarre precisely because of this trait! Her mind is obviously broken in a disturbing way. AIs give me the same feeling. Like talking to an animatronic robot at a theme park or an NPC in a computer game.

Use the memory feature or open the same chat session as before.

Re: ChatGPT-4o vs. Math

#144
post #100

Earlier quoted context omitted.

I am deeply interested in this point of view of yours so I will be hijacking your reply to ask another question: is "better than asking a few random people on the street" the bar we should be setting? As far as mathematical thinking goes this doesn't seem an interesting metric at all. Do you believe that optimizing for this metric will indeed lead to reliable mathematical thinking? I am of the idea that LLMs are not…

People compare a general intelligence against the yardstick of their own specialist skills. I’ve seen some truly absurd examples, like people complaining that it didn’t have the latest updates to some obscure research functional logic proof language that has maybe a hundred users globally! GPT 4 already has markedly superior English comprehension and basic logic than most people I interact with on a daily basis. It’s…

[deleted]

Re: ChatGPT-4o vs. Math

#145
post #96
post #21

I posted the same 'Zero-Shot Chain-of-Thought and Image' to ChatGPT-4o and it made the same error. I then followed up with 'Your math is good but you derived incorrect data from the image. Can you take another look and see if you can tell where the error is?'. It figured it out and corrected it: Let's re-examine the image and the data provided: * The inner radius r1 is given as 5cm * The outer radius r2 is given as 1…

Once you correct the LLM, it will continue to provide the corrected answer until some time later, when it will again make the same mistake. At least, this has been my experience. If you are using LLM to pull answers programmatically and rely on their accuracy, here is what worked for the structured or numeric answers, such as numbers, JSON, etc. 1) Send the same prompt twice, including "Can you double check?" in the…

I can't wait for the day when instead of engineering disciplines solving problems with knowledge and logic they're instead focused on AI/LLM psychology and the correct rituals and incantations that are needed to make the immensely powerful machines at our disposal actually do what we've asked for. /s

Re: ChatGPT-4o vs. Math

#146
post #79

I have a theory that the more you use ChatGPT, the worse it becomes due to silent rate limiting - farming the work out to smaller quantized versions if you ask it a lot of questions. I’d like to see if the results of these tests are the same if you only ask one question per day.

I don't know if that's true, necessarily, but I will note at least anecdotally I find that the larger my context window becomes the more often it seems to make mistakes. eventually, I just have to completely start a new chat even if it's the same topic and thread of conversation.

Re: ChatGPT-4o vs. Math

#147
post #96

Earlier quoted context omitted.

Once you correct the LLM, it will continue to provide the corrected answer until some time later, when it will again make the same mistake. At least, this has been my experience. If you are using LLM to pull answers programmatically and rely on their accuracy, here is what worked for the structured or numeric answers, such as numbers, JSON, etc. 1) Send the same prompt twice, including "Can you double check?" in the…

I can't wait for the day when instead of engineering disciplines solving problems with knowledge and logic they're instead focused on AI/LLM psychology and the correct rituals and incantations that are needed to make the immensely powerful machines at our disposal actually do what we've asked for. /s

"No dude, the bribe you offered was too much so the LLM got spooked, you need to stay in a realistic range. We've fine-tuned a local model on realistic bribe amounts sourced via Mechanical Turk to get a good starting point and then used RLMF to dial in the optimal amount by measuring task performance relative to bribe."

Re: ChatGPT-4o vs. Math

#148
post #143

Earlier quoted context omitted.

Currently, they’re like Dory from Finding Nemo: long and short term memory but they forget everything after each conversation. The character of Dory is jarring and bizarre precisely because of this trait! Her mind is obviously broken in a disturbing way. AIs give me the same feeling. Like talking to an animatronic robot at a theme park or an NPC in a computer game.

Use the memory feature or open the same chat session as before.

Great, so instead of conversing with someone new, we're now conversing with Clive Weaving?

Re: ChatGPT-4o vs. Math

#149
post #113

Earlier quoted context omitted.

A smart human can write and iterate on long, complex chains of logic. We can reason about code bases that are thousands of lines long.

But is that really logic? For instance, we supposedly reason about complex driving laws, but for anyone who has run a stop light late at night when there is no other traffic is acting statistically, not logically.

There's a difference between statistics informing logical reasoning and statistics being used as a replacement for logic.

Running a red light can be perfectly logical. In the mathematics of logic there is no rule that you must obey the law. It can be a calculated risk.

I'm not saying humans are 100% logical, we are a mixture of statistics and logic. What I'm talking about is what we are capable of VS what LLM's are capable of.

I'll give an example. Let's say you give me two random numbers. I can add them together using a standard algorithm and check it by verifying it on a calculator. Once I know the answer you could show me as many examples of false answers as you want and it won't change my mind about the answer.

In LLMs there is clear evidence that the only reason it gets right answers is those answers happen to be more frequent in the dataset. Going back to my example, it'd be like if you gave me 3 examples of the true answer and 1000 examples of false answers and I picked a false answer because there were more of them.

Re: ChatGPT-4o vs. Math

#150

Earlier quoted context omitted.

It's considered an emergent phenomenon of LLMs [1]. So arithmetic reasoning seems to increase as LLMs reasoning grows too. I seem to recall a paper mentioning that LLMs that are better at numeric reasoning are better at overall conversational reasoning too, so it seems like the two come hand in hand. However we don't know the internals of ChatGPT-4, so they may be using some agents to improve performance, or fine-tun…

At the same time the ChatGPT app has access to write and run python, which the gpt can choose to do when it thinks it needs more accuracy.

The results from playing with this are really bizarre: (sorry, formatting hacked up a bit)

To calculate 7^1.83 , you can use a scientific calculator or an exponentiation function in programming or math software. Here is the step-by-step calculation using a scientific calculator:

Input the base: 7 Use the exponentiation function (usually labeled as ^ or x^y). Input the exponent: 1.83 Compute the result. Using these steps, you get:

7^1.83 ≈ 57.864

So, 7^1.83 ≈ 57.864

Given this, and the recent announcement of data analysis features, I’m guessing the GPT-4o is wired up to use various tools, one of which is a calculator. Except that, if you ask it, it also blatantly lies about how it’s using a calculator, and it also sometimes makes up answers (e.g. 57.864 — that’s off by quite a bit).

I imagine some trickery in which the LLM has been trained to output math in some format that the front end can pretty-print, but that there’s an intermediate system that tries (and doesn’t always succeed) to recognize things like “expression =” and emits the tokens for the correct value into the response stream. When it works, great — the LLM magically has correct arithmetic in its output! And when it fails, the LLM cheerfully hallucinates.

Post reply on HN