Earlier quoted context omitted.
> as was requested. They requested week after?
He claims nobody made that request to OpenAI. It was a request made to Google and others who were being judged by the actual contest judges, which OpenAI was not.
Gemini with Deep Think achieves gold-medal standard at the IMO
111–120 of 254 posts
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#112What comes next for this particular exercise? Thank you.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#113Earlier quoted context omitted.
This is a fair reply, but TBH I don't think it's going to change much. The upper echelon of the human society has decided to move AI forward rapidly regardless of any consequences. The rest of us can only hold and pray.
You are watching american money hard at work, my friend. It's either glorious or reckless, hard to tell for now.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#114Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#115I know o3 is far from state of the art these days but it's great at finding relevant literature and suggesting inequalities to consider but in actual proofs it can produce convincing looking statements that are false if you follow the details, or even just the algebra, carefully. Subtle errors like these might become harder to detect as the models get better.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#116Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#117> Btw as an aside, we didn’t announce on Friday because we respected the IMO Board's original request that all AI labs share their results only after the official results had been verified by independent experts & the students had rightly received the acclamation they deserved > We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results off…
I think this is them not being confident enough before the event, so they don't wanna be shown a worse result than competitors. By being private they can obviously not publish anything if it didn't work out.
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#118Solved 5 problems out of 6, scoring 35 out of 42. For comparison OpenAI scored 35/42 too, days back.
Very in character for them!
Re: Gemini with Deep Think achieves gold-medal standard at the IMO
#119Earlier quoted context omitted.
>it seems that the answer to whether or not a general model could perform such a feat is that the models were trained specifically on IMO problems, which is what a number of folks expected. Not sure thats exactly what that means. Its already likely the case that these models contained IMO problems and solutions from pretraining. It's possible this means they were present in the system prompt or something similar.
Does the IMO reuse problems? My understanding is that new problems are submitted each year and 6 are selected for each competition. The submitted problems are then published after the IMO has concluded. How would the training data contain unpublished, newly submitted problems? Obviously the training data contained similar problems, because that's what every IMO participant already studies. It seems unlikely that they…
I don't think I or op suggested it did.