Live data from Hacker News

$10M AI Mathematical Olympiad Prize

aimoprize.com

131–140 of 231 posts

Re: $10M AI Mathematical Olympiad Prize

#131
post #68

I'm asking this question out of ignorance: if you were able to do this, why would you make it public for $10MM instead of keeping it private and exploiting it. Say, in algorithmic trading models?

I see no reason to believe such a machine would be helpful for algorithmic trading. Why should it? How well do current gold medalists do in trading?

Alameda Research had some gold medalists iirc!

Re: $10M AI Mathematical Olympiad Prize

#132

Earlier quoted context omitted.

Well math solving is exactly what the rumored Q* is aiming towards too. I don't think it'll take more than 2 years before some LLM + RL system can take the gold medal. I think companies like OpenAI are aiming for something far more ambitious, like solving a millennium prize problem (even with human assistance). That's the kind of news release that'll add another $100 billion to your market cap.

I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

Yeah, millennium problems almost certainly require truly novel nontrivial ideas to solve.

That's a tough thing for AI to do.

On the other hand, Terrence Tao had an interesting article on his blog a while back where he was trying to solve a problem and asked chatGPT about it in a high-level strategy sense. ChatGPT suggested several reasonable approaches, one of which turned out to work.

That's nowhere near solving a millennium problem, but it is very interesting and suggests fairly sophisticated conceptual understanding of mathematics nevertheless.

Current architecture and training methods I don't think are enough to get there. However, with enough compute, I can plausibly envision some sort of meta training of LLMs using an analogy to GANs where one network tries to synthesize new correct ideas and the other shoots them down as not novel, not correct, or not sufficiently interesting.

Such an approach I think could perhaps work, but the compute needed would probably be pretty high.

Re: $10M AI Mathematical Olympiad Prize

#133
post #86

I was recently in Palo Alto, and bumped into a newly founded startup (I don't remember the name unfortunately) who set themselves the grand the vision of exactly this: winning a gold medal on the international Olympiad using AI. Their plan was to build mostly on LLMs as a start, and iterate as they go. In their barebones office space, they had a poster with a countdown of the number of weeks till the event: it was 36…

I wish there was a serious study of Dunning-Kruger effect in Silicon Valley.

[deleted]

Re: $10M AI Mathematical Olympiad Prize

#134
post #97

Earlier quoted context omitted.

Just to nitpick, while unlikely, it is theoretically possible that the second player has a winning strategy

I would assume this is possible for any sufficiently complex game. Would you mind answering a few questions from someone near-completely ignorant about Go? Does second mover in Go have some sort of artificial benefit in scoring or playing? As in -- is there something to compensate for moving second? On the face it seems like first mover would have an advantage in any turn-based game. But maybe, in some games, seeing…

> Does second mover in Go have some sort of artificial benefit in scoring or playing? As in -- is there something to compensate for moving second?

Yes. The second player typically gets an extra 6.5 or 7.5 points.

Re: $10M AI Mathematical Olympiad Prize

#135

Earlier quoted context omitted.

Well math solving is exactly what the rumored Q* is aiming towards too. I don't think it'll take more than 2 years before some LLM + RL system can take the gold medal. I think companies like OpenAI are aiming for something far more ambitious, like solving a millennium prize problem (even with human assistance). That's the kind of news release that'll add another $100 billion to your market cap.

I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

The field of LLM reasoning is far from stable atm. I think it's pretty hard for anyone to give confident predictions about what they can or cannot do in 5 years. In that light, I am skeptical about any claims that they cannot do something anytime soon.

Re: $10M AI Mathematical Olympiad Prize

#136

Earlier quoted context omitted.

I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

Yeah, millennium problems almost certainly require truly novel nontrivial ideas to solve. That's a tough thing for AI to do. On the other hand, Terrence Tao had an interesting article on his blog a while back where he was trying to solve a problem and asked chatGPT about it in a high-level strategy sense. ChatGPT suggested several reasonable approaches, one of which turned out to work. That's nowhere near solving a m…

And also if you could combine that with some kind of representation software like Lean that can validate proofs, perhaps you can generate some kind of targeted search of the problem space and with a combination of a general high level strategy and maybe brute force of some sub-problems, find novel solutions. My understanding is that is the common human approach: gain some kind of intuition of problem, try a few things and then iteratively refine. Sometimes that works, sometimes you need to find a new starting point. It seems plausible we could automate that workflow, with likely mixed but still useful results.

Re: $10M AI Mathematical Olympiad Prize

#137

Earlier quoted context omitted.

I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

Yeah, millennium problems almost certainly require truly novel nontrivial ideas to solve. That's a tough thing for AI to do. On the other hand, Terrence Tao had an interesting article on his blog a while back where he was trying to solve a problem and asked chatGPT about it in a high-level strategy sense. ChatGPT suggested several reasonable approaches, one of which turned out to work. That's nowhere near solving a m…

I’m not trying to understate the stuff you can do with LLMs or formal proof tooling that could be attached to randomly try things till it reaches a solution to a novel problem, but some solutions you see to lesser problems are sometimes so pie-in-the-sky I think stumbling on one is barely better than a random walk. And I think mathematicians like Tao are far better at narrowing that down. As an aide I can see it’s use, I just don’t believe we’re going to have LLMs & tooling solve these grand problems until compute power is orders of magnitudes better, and even then I’m still sceptical. BUT I’m not a mathematician and this is based on my intuition from pub talks :)

Re: $10M AI Mathematical Olympiad Prize

#138
post #29

Earlier quoted context omitted.

Their current ambition is to be able to solve school math, which is quite far away from solving unsolved conjectures or math olympiads. I really doubt that any of this is within LLM/transformer scope, except maybe in some auxiliary sense to other, much different architectures.

Art isn't an easier problem than math. An artbot would have sounded more sci-fi than a mathbot only 2 years ago. Yet it only took the AI world 1.5 years to go from drawing child scribbles to replicating top artists with like 90% similarity (I can barely tell the difference between AI and human drawn art anymore with the new NovelAI model). It won't be long before AI starts to go superhuman in art skills. It won't tak…

> Art isn't an easier problem than math.

Someone can tell you when a math problem is solved. Someone else can't tell you when you've successfully art'd with remotely the same degree of confidence. In so far as they can, however, many experts claim that AI cannot, indeed "do an art".

Also, it's easier to formulate a hard math question (there are plenty of unsolved problems), but that's (IMHO) harder to do for art. Sure, you may think this is the first time the phrase "Astronaut riding a Llama and holding an avocado" was writ, but those are all well represented concepts in the dataset. For more abstract prompts, there really isn't a way to verify "correctness".

Re: $10M AI Mathematical Olympiad Prize

#139
post #86

I was recently in Palo Alto, and bumped into a newly founded startup (I don't remember the name unfortunately) who set themselves the grand the vision of exactly this: winning a gold medal on the international Olympiad using AI. Their plan was to build mostly on LLMs as a start, and iterate as they go. In their barebones office space, they had a poster with a countdown of the number of weeks till the event: it was 36…

I wish there was a serious study of Dunning-Kruger effect in Silicon Valley.

While I get your point and agree, there apparently is no Dunning-Kruger effect. It was determined to be yet another case of faulty data analysis in behavioral psych.[1]

1. https://economicsfromthetopdown.com/2022/04/08/the-dunning-k...

Re: $10M AI Mathematical Olympiad Prize

#140

Earlier quoted context omitted.

I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

The field of LLM reasoning is far from stable atm. I think it's pretty hard for anyone to give confident predictions about what they can or cannot do in 5 years. In that light, I am skeptical about any claims that they cannot do something anytime soon.

What does an LLM really do? And can it create sometimes entirely novel math that goes against its training set to solve something, and know it’s right using a proof tool that may not even accept that? I agree but on a different order of magnitude of years.
Post reply on HN