Live data from Hacker News

ChatGPT-4o vs. Math

sabrina.dev

31–40 of 182 posts

Re: ChatGPT-4o vs. Math

#31

Similar to the article, I haven't found complementary image data to be that useful. If the information is really missing without the image, then the image is useful. But if the basic information is all available textually (including things like the code that produces a diagram) then the image doesn't seem to add much except perhaps some chaos/unpredictability. But reading this I do have a thought: chain of thought, o…

Yeah, like the other commenter mentioned, I could have run another experiment applying chain of thought specifically to the image interpretation. Just to force gpt to confirm its information extraction from the image. However, even after trying that approach, it got only 2/3 tries correct. Still superior is text only modality + chain of thought.

Re: ChatGPT-4o vs. Math

#32
post #3

The model's first attempt is impressive (not sure why it's labeled a choke). Unfortunately gpt4o cannot discover calculus on its own.

I don't know... here's a prompt query for a standard problem in introductory integral calculus, and it seems to go pretty smoothly from a discrete arithmetical series into the continuous integral:

"Consider the following word problem: "A 100 meter long chain is hanging off the end of a cliff. It weighs one metric ton. How much physical work is required to pull the chain to the top of the cliff if we discretize the problem such that one meter is pulled up at a time?" Note that the remaining chain gets lighter after each lifting step. Find the equation that describes this discrete problem and from that, generate the continuous expression and provide the Latex code for it."

Re: ChatGPT-4o vs. Math

#33
post #21

I posted the same 'Zero-Shot Chain-of-Thought and Image' to ChatGPT-4o and it made the same error. I then followed up with 'Your math is good but you derived incorrect data from the image. Can you take another look and see if you can tell where the error is?'. It figured it out and corrected it: Let's re-examine the image and the data provided: * The inner radius r1 is given as 5cm * The outer radius r2 is given as 1…

This speaks to a deeper issue that LLMs don’t just have statistically-based knowledge, they also have statistically-based reasoning.

This means their reasoning process isn’t necessarily based on logic, but what is statistically most probable. As you’ve experienced, their reasoning breaks down in less-common scenarios even if it should be easy to use logic to get the answer.

Re: ChatGPT-4o vs. Math

#34
post #19

Earlier quoted context omitted.

A lot of credible people have left Twitter - it has gotten much more overrun by bots and a lot of very hateful accounts have been reinstated and protected. It is a poor platform for reasonable discussion and I think it's fair to say it's been stifling open expression. The value is disappearing.

that was not the question

I think it was an appropriate answer at the heart of the matter - most credible people are leaving the platform due to the degradation of quality on it. For a literal example of a ban though there are few examples better than Dell Cameron[1].

1. https://www.vanityfair.com/news/2023/04/elon-musk-twitter-st...

Re: ChatGPT-4o vs. Math

#35
post #21

I posted the same 'Zero-Shot Chain-of-Thought and Image' to ChatGPT-4o and it made the same error. I then followed up with 'Your math is good but you derived incorrect data from the image. Can you take another look and see if you can tell where the error is?'. It figured it out and corrected it: Let's re-examine the image and the data provided: * The inner radius r1 is given as 5cm * The outer radius r2 is given as 1…

This speaks to a deeper issue that LLMs don’t just have statistically-based knowledge, they also have statistically-based reasoning. This means their reasoning process isn’t necessarily based on logic, but what is statistically most probable. As you’ve experienced, their reasoning breaks down in less-common scenarios even if it should be easy to use logic to get the answer.

Does anyone know how far off we are having logical AI?

Math seems like low hanging fruit in that regard.

But logic as it's used in philosophy feels like it might be a whole different and more difficult beast to tackle.

I wonder if LLM's will just get better to the point of being indistinguishable from logic rather than actually achieving logical reasoning.

Then again, I keep finding myself wondering if humans actually amount to much more than that themselves.

Re: ChatGPT-4o vs. Math

#36

Earlier quoted context omitted.

This speaks to a deeper issue that LLMs don’t just have statistically-based knowledge, they also have statistically-based reasoning. This means their reasoning process isn’t necessarily based on logic, but what is statistically most probable. As you’ve experienced, their reasoning breaks down in less-common scenarios even if it should be easy to use logic to get the answer.

Does anyone know how far off we are having logical AI? Math seems like low hanging fruit in that regard. But logic as it's used in philosophy feels like it might be a whole different and more difficult beast to tackle. I wonder if LLM's will just get better to the point of being indistinguishable from logic rather than actually achieving logical reasoning. Then again, I keep finding myself wondering if humans actuall…

(Not an AI researcher, just someone who likes complexity analysis.) Discrete reasoning is NP-Complete. You can get very close with the stats-based approaches of LLMs and whatnot, but your minima/maxima may always turn out to be local rather than global.

Re: ChatGPT-4o vs. Math

#37
post #15

Earlier quoted context omitted.

I mean, in this context I agree. But most people doing math in high school or university are graded on their working of a problem, with the final result usually equating to a small proportion of the total marks received.

But most people doing math in high school or university are graded on their working of a problem, with the final result usually equating to a small proportion of the total marks received That heavily depends on the individual grader/instructor. A good grader will take into account the amount of progress toward the solution. Restating trivial facts of the problem (in slightly different ways) or pursuing an invalid sol…

it choked because it didn't solve for `t` at the end

impressive attempt though, it used number of wraps which I found quite clever

Re: ChatGPT-4o vs. Math

#38
This recent article on Hacker News seems to suggest similar inconsistencies.

GPT-4 Turbo with Vision is a step backward for coding (aider.chat) https://news.ycombinator.com/item?id=39985596

Without looking deeply at how cross-attention works, I imagine the instruction tuning of the multimodal models to be challenging.

Maybe the magic is in synthetically creating this instruct dataset that combines images and text in all the ways they can relate. I don't know if I can even begin to imagine how they could be used together.

Re: ChatGPT-4o vs. Math

#39
post #2

Need to run this experiment on a problem that is t already on its training set.

Is there any good literature on this topic?

I feel like math is naturally one of the easiest sets of synthetic data we can produce, especially since you can represent the same questions multiple ways in word problems.

You could just increment the numbers infinitely and generate billions of examples of every formula.

If we can't train them to be excellent at math, what hope do we ever have at programming or any other skill?

Re: ChatGPT-4o vs. Math

#40
post #19

Earlier quoted context omitted.

A lot of credible people have left Twitter - it has gotten much more overrun by bots and a lot of very hateful accounts have been reinstated and protected. It is a poor platform for reasonable discussion and I think it's fair to say it's been stifling open expression. The value is disappearing.

that was not the question

[flagged]
Post reply on HN