Live data from Hacker News

ChatGPT-4o vs. Math

sabrina.dev

11–20 of 182 posts

Re: ChatGPT-4o vs. Math

#12
post #4

Earlier quoted context omitted.

It's a choke because it failed to get the answer. Saying other true things but not getting the answer is not a success.

I mean, in this context I agree. But most people doing math in high school or university are graded on their working of a problem, with the final result usually equating to a small proportion of the total marks received.

This is supposed to be a product , not a research artifact.

Re: ChatGPT-4o vs. Math

#14
post #10

As an aside, what did the author do to get banned on X?

Not gonna lie, when I see that someone is banned on X, I assume credibility

Even the 100's of Hamas-affiliated accounts?

https://ny1.com/nyc/all-boroughs/technology/2023/10/12/x-say...

Re: ChatGPT-4o vs. Math

#15
post #4

Earlier quoted context omitted.

It's a choke because it failed to get the answer. Saying other true things but not getting the answer is not a success.

I mean, in this context I agree. But most people doing math in high school or university are graded on their working of a problem, with the final result usually equating to a small proportion of the total marks received.

But most people doing math in high school or university are graded on their working of a problem, with the final result usually equating to a small proportion of the total marks received

That heavily depends on the individual grader/instructor. A good grader will take into account the amount of progress toward the solution. Restating trivial facts of the problem (in slightly different ways) or pursuing an invalid solution to a dead end should not be awarded any marks.

Re: ChatGPT-4o vs. Math

#16
post #10

Earlier quoted context omitted.

Not gonna lie, when I see that someone is banned on X, I assume credibility

Even the 100's of Hamas-affiliated accounts? https://ny1.com/nyc/all-boroughs/technology/2023/10/12/x-say...

straw man, and a drop in the bucket

Re: ChatGPT-4o vs. Math

#18
post #3

The model's first attempt is impressive (not sure why it's labeled a choke). Unfortunately gpt4o cannot discover calculus on its own.

I think this is the biggest flaw in LLMs and what is likely going to sour a lot of businesses on their usage (at least in their current state). It is preferable to give the right answer to a query, it is acceptable to be unable to answer a query - we run into real issues, though, when a query is confidently answered incorrectly. This recently caused a major headache for AirCanada - businesses should be held to the statements they make, even if those statements were made by an AI or call center employee.

Re: ChatGPT-4o vs. Math

#19
post #10

Earlier quoted context omitted.

Not gonna lie, when I see that someone is banned on X, I assume credibility

Why? Are many credible people banned on Twitter?

A lot of credible people have left Twitter - it has gotten much more overrun by bots and a lot of very hateful accounts have been reinstated and protected. It is a poor platform for reasonable discussion and I think it's fair to say it's been stifling open expression. The value is disappearing.

Re: ChatGPT-4o vs. Math

#20
post #19

Earlier quoted context omitted.

Why? Are many credible people banned on Twitter?

A lot of credible people have left Twitter - it has gotten much more overrun by bots and a lot of very hateful accounts have been reinstated and protected. It is a poor platform for reasonable discussion and I think it's fair to say it's been stifling open expression. The value is disappearing.

that was not the question
Post reply on HN