Which is greater, 9.11 or 9.9?
/sI kid, this is actually pretty amazing!! I've noticed over the last several months that I've had to correct it less and less when dealing with advanced math topics so this aligns.
51–60 of 737 posts
Which is greater, 9.11 or 9.9?
/sI kid, this is actually pretty amazing!! I've noticed over the last several months that I've had to correct it less and less when dealing with advanced math topics so this aligns.
Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…
Earlier quoted context omitted.
Nope https://x.com/polynoamial/status/1946478249187377206?s=46&t=...
If you don't have a Twitter account then x.com links are useless, use a mirror: https://xcancel.com/polynoamial/status/1946478249187377206 Anyway, that doesn't refute my point, it's just PR from a weaselly and dishonest company. I didn't say it was "IMO-specific" but the output strongly suggests specialized tooling and training, and they said this was an experimental LLM that wouldn't be released. I strongly suspect…
Wow. That's an impressive result, but how did they do it? Wei references scaling up test-time compute, so I have to assume they threw a boatload of money at this. I've heard talk of running models in parallel and comparing results - if OpenAI ran this 10000 times in parallel and cherry-picked the best one, this is a lot less exciting. If this is legit, then we need to know what tools were used and how the model used…
According to the twitter thread, the model was not given access to tools.
Earlier quoted context omitted.
> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.
[flagged]
Many of them are also questions that eventually end up with proofs or solutions that only require very high level understanding of basic principles. But when I say very high I mean like impossibly high for the average person and ability to combine simple concepts to solve complex problems.
I'd wager the majority of Math graduates from universities would struggle to answer most IMO questions.
Earlier quoted context omitted.
> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.
[flagged]
Earlier quoted context omitted.
Getting gold at the IMO is pretty damn hard. I grew up in a relatively underserved rural city. I skipped multiple grades in math, completed the first two years of college math classes while in high school, and won the award for being the best at math out of everyone in my school. I've met and worked with a few IMO gold medalists. Even though I was used to scoring in the 99th percentile on all my tests, it felt like t…
The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)
Earlier quoted context omitted.
The trouble is, getting an IMO gold medal is much easier (by frequency) than being the #1 Go player in the world, which was achieved by AI 10 years ago. I'm not sure it's enough to just gesture at the task; drilling down into precisely how it was achieved feels important. (Not to take away from the result, which I'm really impressed by!)
The "AI" that won Go was Monte Carlo tree search on a neural net "memory" of the outcome of millions of previous games; this is a LLM solving open ended problems. The tasks are hardly even comparable.
https://deepmind.google/discover/blog/alphago-zero-starting-...
Earlier quoted context omitted.
> high school/early university maths problems should not have been a stretch at all for it This is a ridiculous understatement of the difficulty of getting gold at the IMO.
[flagged]