It makes too many mistakes and is just way too sloppy with math. It shouldn't be this hard to do pair-theorem-proving with it. It cannot tell the difference between a conjecture that sounds kind of vaguely plausible and something that is actually true, and literally the entire point of math is to successfully differentiate between those two situations. It needs to be able to carefully keep track of which claims it's…
Curious to know how the different models compare for you for doing math. Heard o4-mini is really good at math but haven’t tried o3-pro much.
o3-pro is maybe marginally better, but it takes a very long time to respond and so I rarely use it.
4o is much worse and so I usually use o3.
Gemini 2.5 Pro is much better - and free. Grok 4 is also probably up there with Gemini 2.5. They just have less tendency to hallucinate in this way in general: they will spend more time reasoning, checking claims, searching for prior literature, etc. They still mess up, but not quite as much as o3. I don't use Sonnet or Opus for math all that much - my impression was that o3 was better than Sonnet 3.7 but not sure about 4.