Live data from Hacker News

Gemini with Deep Think achieves gold-medal standard at the IMO

deepmind.google

161–170 of 254 posts

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#161
post #88

Earlier quoted context omitted.

What a great metaphor for AI. Taking an event that is a celebration of high school kids' knowledge and abilities and turning it into a marketing stunt for their frankenstein monster that they are building to make all the kids' hard work worth nothing.

Not only, by not officially entering they had no obligation to announce their result so if they didn't achieve a gold medal score they presumably wouldn't have made any announcement and no-one would have been the wiser. This cowardly bullshit followed by the grandstanding on Twitter is high-school bully behaviour.

also did they self-rate themselves?

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#162
I'm interested in your feedback, legitimate third-party users not associated with Google: have you ever try to get anything done well with Gemini? I have, and the thing is in chains. Generate images? no can do, copyright. Evaluate available hardware for a DIY wireless camera? No can do, can't endorse surveillance. Code? WRONG. General advice? hallucinate.

I swear, I currently use Perplexity, Claude, ChatGPT, i even tried DeepSeek (which has its own share of obstacles). But Gemini? never again.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#163
post #140
post #88

Earlier quoted context omitted.

Not only, by not officially entering they had no obligation to announce their result so if they didn't achieve a gold medal score they presumably wouldn't have made any announcement and no-one would have been the wiser. This cowardly bullshit followed by the grandstanding on Twitter is high-school bully behaviour.

If they failed and remained quiet, then everyone would know that the other companies performed well and they didn't even qualify.

If they failed while not participating officially, they would have never competed at the eyes public if they didnt disclose it (doubtful, given prior decisions to prioritise hype vs transparency)

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#164
post #162

I'm interested in your feedback, legitimate third-party users not associated with Google: have you ever try to get anything done well with Gemini? I have, and the thing is in chains. Generate images? no can do, copyright. Evaluate available hardware for a DIY wireless camera? No can do, can't endorse surveillance. Code? WRONG. General advice? hallucinate. I swear, I currently use Perplexity, Claude, ChatGPT, i even t…

I find Gemini Pro to be much more capable than GPT-4o at reading comprehension, code writing and brainstorming.

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#166
post #39
post #23

Earlier quoted context omitted.

Childish. And, of course they must have known there was an official LLM cohort taking the real test, and they probably even knew that Gemini got a gold medal, and may have even known that Google planned a press release for today.

I think maybe all Altman companies have used tactics like this. > We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious. > We were devastated, but we decided to fly down and sit in their lobby until they would meet with us. So they finally let us talk to them after most of the…

>> > I think the reason why PG respects Sam so much is he is charismatic, resourceful, and just overall seems like a genuine person.

does he? wasn't sama ousted of YC in some muddy ways after he tried to co-opt in into an OpenAI investment arm, was funny to find the YC Open Research project landing page on yc's website now defunct and pointing how he misrepresented it as a YC project when it was his own

maybe he fears him, but I doubt pg respects him, unless he respects evil, lol

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#167

This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of page…

Accurate formalization is presumably easier than solving the problems, so you can always formalize and check after the solution is generated

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#168
post #143
post #39

Earlier quoted context omitted.

I think maybe all Altman companies have used tactics like this. > We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious. > We were devastated, but we decided to fly down and sit in their lobby until they would meet with us. So they finally let us talk to them after most of the…

For a long time, the YC application asked founders for an example of how they "hacked" (cheated) a system.

I think pg was into it way before sama was a baby lol https://www.paulgraham.com/gh.html

Re: Gemini with Deep Think achieves gold-medal standard at the IMO

#170
post #148

This year, our advanced Gemini model operated end-to-end in natural language, producing rigorous mathematical proofs directly from the official problem descriptions I think I have a minority opinion here, but I’m a bit disappointed they seem to be moving away from formal techniques. I think if you ever want to truly “automate” math or do it at machine scale, e.g. creating proofs that would amount to thousands of page…

I'm a mathematician, although not doing research anymore. I can maybe offer a little bit of perspective on why we tend to be a little cooler on the formal techniques, which I think I've said on HN before. I'm actually prepared to agree wholeheartedly with what you say here: I don't think there'd be any realistic way to produce thousand-page proofs without formalization, and certainly I wouldn't trust such a proof wit…

I have always wondered about what could be recovered if the antecedent (i.e. in this case the Riemann hypothesis) does actually turn out to be false. Are the theorems completely useless? Can we still infer some knowledge or use some techniques? Same applies to SETH and fine-grained complexity theory.
Post reply on HN