Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

101–110 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#101

I think we are having a Deep Blue vs. Kasparov moment in Competitive Math right now. This is a large progress from just a few years ago and yet I think we still are really far away from even a semi-respectable AI mathematician. What an exciting time to be alive!

Terrence Tao, in a recent podcast, said that he's very interested in "working along side these tools". He sees the best use in the near future as "explorers of human set vision" in a way. (i.e. set some ideas/parameters and let the LLMs explore and do parallel search / proof / etc) Your comparison with chess engines is pretty spot-on, that's how the best of the best chess players do prep nowadays. Gone are the multi…

> They now have analysts that use supercomputers to search through bajillions of positions and extract the best ideas, and distill

I was recently researching AI's for this, seems it would be a huge unlock for some parts of science where this is the case too like chess

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#102

From Terence Tao, via mastodon [0]: > It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. > One can illustrate this with a hum…

This is a fair reply, but TBH I don't think it's going to change much. The upper echelon of the human society has decided to move AI forward rapidly regardless of any consequences. The rest of us can only hold and pray.

You are watching american money hard at work, my friend. It's either glorious or reckless, hard to tell for now.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#103

This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of page…

(Stream of consciousness aside:

That said, letting machines go wild in the depths of the consequences of some axiomatic system like ZFC may reveal a method of proof mathematicians would find to be monstrous. So like, if ZFC is inconsistent, then anything can be proven. But short of that, maybe the machines will find extremely powerful techniques which “almost” prove inconsistency that nevertheless somehow lead to logical proofs of the desired claim. I’m thinking by analogy here about how speedrunning seems to often devolve into exploiting an ACE glitch as early as possible, thus meeting the logical requirements of finishing the game while violating the spirit. Maybe we’d have to figure out what “glitchless ZFC” should mean. Maybe this is what logicians have already been doing heh).

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#104

Earlier quoted context omitted.

Sounds like it did not: > This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions – all within the 4.5-hour competition time limit

Yes, that quote is contained in my comment. But I don't think it unambiguously excludes tool use in the internal chain of thought. I don't think tool use would detract from the achievement, necessarily. I'm just interested to know.

End to end in natural language would imply no tool use, I'd imagine. Unless it called another tool which converted it but that would be a real stretch (smoke and mirrors).

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#105
post #39
post #23

Earlier quoted context omitted.

Childish. And, of course they must have known there was an official LLM cohort taking the real test, and they probably even knew that Gemini got a gold medal, and may have even known that Google planned a press release for today.

I think maybe all Altman companies have used tactics like this. > We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious. > We were devastated, but we decided to fly down and sit in their lobby until they would meet with us. So they finally let us talk to them after most of the…

So, a more charismatic version of Zuck is Zucking, what a surprise. Company culture starts at its origin. Despite Google's corruption, its origin is in academia and it shows even now.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#106
post #89

Earlier quoted context omitted.

That PDF lists solutions for problems 1 through 5 but does not mention problem 6 at all.

Ah thanks that does answer it. You should start a blog

Simon's a prolific writer on AI

http://simonw.substack.com https://simonwillison.net

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#108
post #56

Surprising since Reid Barton was working on a lean system.

The bitter lesson.

well lean systems might be still useful for other stuff than max benching

my point being transformers and llms have all the tailwind of all the infra+lateral discoveries/improvements being put into them.

does that mean they're the one tool to unlock machine intelligence? I dunno

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#109
post #61

Earlier quoted context omitted.

They will have to rename gemini anyway, since it's doubtful they will ever be able to buy gemini.com. Gemini turboflex plus pro SE plaid interstellar 4.0

What makes you think google can't buy that domain?

... for a reasonable cost.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#110

This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of page…

These problems are designed to be solvable by humans without tools. No reason we can't give tools to the models when they go after harder problems. I think it's good to at least reproduce human-level skill without tools first.
Post reply on HN