Live data from Hacker News

ChatGPT-4o vs. Math

sabrina.dev

121–130 of 182 posts

Re: ChatGPT-4o vs. Math

#121
post #56
post #43

Earlier quoted context omitted.

I think LLMs will need to do what humans do: invent symbolic representations of systems and then "reason" by manipulating those systems according to rules. Here's a paper working along those lines: https://arxiv.org/abs/2402.03620

Is this what humans do?

No. Not in my experience. Anyone with experience in research mathematics will tell you that making progress at the research level is driven by intuition - intuition honed from years of training with formal rules and rigor but intuition nonetheless - with the final step being to reframe the argument in formal/rigorous language and ensure consitency and so forth.

Infact the more experience and skill I get in supposedly "rational" subjects like foundations, set theory, theoretical physics, etc. the more sure I am that intuition / belief first - justification later is a fundamental tenant of how human brains operate, and the key feature of rationalism and science during the enlightenment was producing a framework so that one may have some way to sort beliefs, theories, and assertion so that we can recover - at the end - some kind of gesture towards objectivity

Re: ChatGPT-4o vs. Math

#122
post #86
post #56

Earlier quoted context omitted.

Is this what humans do?

Think of all the algebra problems you got in school where the solution started with "get all the x's on the same side of the equation." You then applied a bunch of rules like "you can do anything to one side of the equals sign if you also do it to the other side" to reiterate the same abstract concept over and over, gradually altering the symbology until you wound up at something that looked like the quadratic formul…

People don't uncover new mathematics with formal rules and symbols pushing, at least not for the most part. They do so first with intuition and vague belief. Formalisation and rigour is the final stage of constructing a proof or argument.

Re: ChatGPT-4o vs. Math

#123

I actually have a contrarian view: being able to do elementary math is not that important in the current stage. Yes, understanding elementary math is a cornerstone for an AI to become more intelligent, but also let's be honest: LLMs are far from being AGIs and does not have common sense nor general ability to deduce or induct. If we accept such limitation of LLM, then focusing the mathematical understanding of an LLM…

It’s important because solving a math problem requires you to actually understand something and follow deliberate steps. The fact that they can’t means they’re just a toy ultimately.

No, I disagree. It is just deliberate steps. Understanding can greatly help you do the steps and remember which ones to do.

Training math is likely hard because the corpus of training data is so much less because the computers themselves do our math as it relates to computers. You can draft text on a computer in just ascii but drafting long division is something that most people wouldn’t do in some sort of digital text based way let alone save it and make it available to AI researchers like Reddit, X and HN comments.

I expect LLMs to be bad at math. That’s ok, they are bad because the computers themselves are so good at math.

Re: ChatGPT-4o vs. Math

#124
post #101

Do we know why GPT-4o seems able to do arithmetic? Is it outsourcing to some tool?

It's considered an emergent phenomenon of LLMs [1]. So arithmetic reasoning seems to increase as LLMs reasoning grows too. I seem to recall a paper mentioning that LLMs that are better at numeric reasoning are better at overall conversational reasoning too, so it seems like the two come hand in hand. However we don't know the internals of ChatGPT-4, so they may be using some agents to improve performance, or fine-tun…

At the same time the ChatGPT app has access to write and run python, which the gpt can choose to do when it thinks it needs more accuracy.

Re: ChatGPT-4o vs. Math

#125
post #51

This problem strikes me as relatively simple. What about more complex math problems? Are there good benchmarks for that? I would dearly love to have an AI tool that I could trust to help with math. What is the state of the art? My math skills are very rusty (the last math class I took was calculus almost 40 years ago), and I find myself wanting to do things which would require a PhD level understanding of computer ai…

ChatGPT has an amazing ability to write, but you shouldn't trust it for any form of mathematics aside from providing vague descriptions of what various topics are about (and even that tends to result in a word soup that is more flowery than descriptive). When it comes to solving specific problems, or even providing specific examples of mathematical objects, it falls down really quickly. I'll inevitably be told otherw…

Terry Tao finds it promising https://mathstodon.xyz/@tao/110601051375142142

I am a first year grad student and find it useful to chat about stuff with Claude, especially once my internal understanding has just gotten clarified. It isn't as good as the professor but is available at 2 am.

Re: ChatGPT-4o vs. Math

#126

Earlier quoted context omitted.

ChatGPT has an amazing ability to write, but you shouldn't trust it for any form of mathematics aside from providing vague descriptions of what various topics are about (and even that tends to result in a word soup that is more flowery than descriptive). When it comes to solving specific problems, or even providing specific examples of mathematical objects, it falls down really quickly. I'll inevitably be told otherw…

> ChatGPT-happy hypebro Rude. From the guidelines: > Please don't sneer, including at the rest of the community. https://news.ycombinator.com/newsguidelines.html "math help" is really broad, but if you add "solve this using python", chatgpt will generate code and run that instead of trying to do logic as a bare LLM. There's no guarantee that it gets the code right, so I won't claim anything about its reliability, but…

You’re right, but I get frustrated by the ignorance and hubris of some people. Too late to edit now.

Re: ChatGPT-4o vs. Math

#127
If you really want to see what the SOTA model can do, look at the posts on the web page for the mind-blowing image output. That is not released yet. https://openai.com/index/hello-gpt-4o/

Mark my words, that is the sort of thing that Ilya saw months ago and I believe he decided they had achieved their mission of AGI. And so that would mean stopping work, giving it to the government to study, or giving it away or something.

That is the reason for the coup attempt. Look at the model training cut-off date. And Altman won because everyone knew they couldn't make money by giving it away if they just declared mission accomplished and gave it away or to some government think-tank and stopped.

This is also why they didn't make a big deal about those capabilities during the presentation. Because if they go too hard on the abilities, more people will start calling it AGI. And AGI basically means the company is a wrap.

Re: ChatGPT-4o vs. Math

#129

Earlier quoted context omitted.

ChatGPT has an amazing ability to write, but you shouldn't trust it for any form of mathematics aside from providing vague descriptions of what various topics are about (and even that tends to result in a word soup that is more flowery than descriptive). When it comes to solving specific problems, or even providing specific examples of mathematical objects, it falls down really quickly. I'll inevitably be told otherw…

Terry Tao finds it promising https://mathstodon.xyz/@tao/110601051375142142 I am a first year grad student and find it useful to chat about stuff with Claude, especially once my internal understanding has just gotten clarified. It isn't as good as the professor but is available at 2 am.

I think Tao finds it promising as a source of inspiration in the same sense that the ripples on the surface of a lake or a short walk in the woods can be mathematically inspiring. It doesn’t say much about the actual content being produced; the more you already have going on in your head the more easily you ascribe meaning to meaninglessness.

The point is that it’s got seemingly nothing to do with reasoning. That it can produce thought-stimulating paragraphs about any given topic doesn’t contradict that; chatting to something not much more sophisticated than Eliza (or even… yourself, in a mirror) could probably produce a similar effect.

As for chatting about stuff, I’ve been experimenting with ChatGPT a bit for that kind of thing but find its output usually too vague. It can’t construct examples of things beyond the trivial/very standard ones that don’t say much, and that’s assuming it’s even getting it right which it often isn’t (it will insist on strange statements despite also admitting them to be false). It’s a good memory-jog for things you’ve half forgotten, but that’s about it.

Re: ChatGPT-4o vs. Math

#130
post #118

LLMs are deterministic with 0 temperature on the same hardware with the same seed though, as long as the implementation is deterministic. You can easily use the OpenAI API with the temp=0 and a predefined seed and you'll get very deterministic results

> You can easily use the OpenAI API with the temp=0 and a predefined seed and you'll get very deterministic results Does that mean that in this situation OpenAI will always answer wrongly for the same question?

temp 0 means that there will be no randomness injected into the response, and that for any given input you will get the exact same output, assuming the context window is also the same. Part of what makes an LLM more of a "thinking machine" than purely a "calculation machine" is that it will occasionally choose a less-probable next token than the statistically most likely token as a way of making the response more "flavorful" (or at least that's my understanding of why), and the likelihood of the response diverging from its most probable outcome is influenced by the temperature.
Post reply on HN