Live data from Hacker News

Gemini 3 Deep Think

blog.google

511–520 of 722 posts

Re: Gemini 3 Deep Think

#512
post #3

Here is the methodologies for all the benchmarks: https://storage.googleapis.com/deepmind-media/gemini/gemini_... The arc-agi-2 score (84.6%) is from the semi-private eval set. If gemini-3-deepthink gets above 85% on the private eval set, it will be considered "solved" >Submit a solution which scores 85% on the ARC-AGI-2 private evaluation set and win $700K. https://arcprize.org/guide#overview

Huh, so if a China-based lab takes ARC-AGI-2 on the new year, then they can say they had just-shy of a solution anyway.

Re: Gemini 3 Deep Think

#513

Earlier quoted context omitted.

Aren't we saying "lunar new year" now?

I don't think so; there are different lunar calendars.

If that's a sole problem, it should be called "Chinese-Japanese-Korean-whateverelse new year" instead. Maybe "East Asian new year" for short. (Not that there are absolutely no discrepancies within them, but they are so similar enough that new year's day almost always coincide.)

Re: Gemini 3 Deep Think

#514

Earlier quoted context omitted.

Ok, here I am living in the real world finding these models have advanced incredibly over the past year for coding. Benchmaxxing exists, but that’s not the only data point. It’s pretty clear that models are improving quickly in many domains in real world usage.

I agree completely. I think we're in alignment with Elon Musk who says that AI will bypass coding entirely and create the binary directly. It's going to be an exciting year.

There’s about as much sense doing this as there is in putting datacenters in orbit, i.e. it isn’t impossible, but literally any other option is better.

Re: Gemini 3 Deep Think

#515

Earlier quoted context omitted.

When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.

Vetting them for the potential for whistleblowing might be a bit more involved. But conspiracy theories have an advantage because the lack of evidence is evidence for the theory.

Huh? AI labs are routinely spending millions to billions to various 3rd party contractors specializing in creating/labeling/verifying specialized content for pre/post-training.

This would just be one more checkbox buried in hundreds of pages of requests, and compared to plenty of other ethical grey areas like copyright laundering with actual legal implications, leaking that someone was asked to create a few dozen pelican images seems like it would be at the very bottom of the list of reputational risks.

Re: Gemini 3 Deep Think

#516

Earlier quoted context omitted.

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

Isn’t that super intelligence not AGI? Feels like these benchmarks continue to move the goalposts.

It's probably both. We've already achieved superintelligence in a few domains. For example protein folding.

AGI without superintelligence is quite difficult to adjudicate because any time it fails at an "easy" task there will be contention about the criteria.

Re: Gemini 3 Deep Think

#517

Earlier quoted context omitted.

These are well behind the general state of the art (1yr or so), though they're arguably the best openly-available models.

Idk man, GLM 5 in my tests matches opus 4.5 which is what, two months old?

4.5 was never sota

Re: Gemini 3 Deep Think

#518

Earlier quoted context omitted.

Strange, because I could not for the life of me get Gemini 3 to follow my instructions the other day to work through an example with a table, Claude got it first try.

Claude is king for agentic workflows right now because it’s amazing at tool calling and following instructions well (among other things)

Codex ranks higher for instruction following

Re: Gemini 3 Deep Think

#519

Earlier quoted context omitted.

> which is why only "ARC-AGI Certified" results using a secret problem set really matter. The 84.6% is certified and that's a pretty big deal. So, I'd agree if this was on the true fully private set, but Google themselves says they test on only the semi-private: > ARC-AGI-2 results are sourced from the ARC Prize website and are ARC Prize Verified. The set reported is v2, semi-private ( https://storage.googleapis.com/…

Chollet himself says "We certified these scores in the past few days." https://x.com/fchollet/status/2021983310541729894 . The ARC-AGI papers claim to show that training on a public or semi-private set of ARC-AGI problems to be of very limited value in passing a private set. none of ARC-AGI can possibly be valid. So, before "public, semi-private or private" answers leaking or 'benchmaxing' on them can even matter - y…

They could also cheat on the private set though. The frontier models presumably never leave the provider's datacenter. So either the frontier models aren't permitted to test on the private set, or the private set gets sent out to the datacenter.

But I think such quibbling largely misses the point. The goal is really just to guarantee that the test isn't unintentionally trained on. For that, semi-private is sufficient.

Re: Gemini 3 Deep Think

#520
post #175

Earlier quoted context omitted.

I think it is because of the Chinese new year. The Chinese labs like to publish their models arround the Chinese new year, and the US labs do not want to let a DeepSeek R1 (20 January 2025) impact event happen again, so i guess they publish models that are more capable then what they imagine Chinese labs are yet capable of producing.

I guess. Deepseek v3 was released on boxing day a month prior https://api-docs.deepseek.com/news/news1226

And made almost zero impact, it was just a bigger version of Deepseek V2 and when mostly unnoticed because its performances weren't particularly notable especially for its size.

It was R1 with its RL-training that made the news and crashed the srock market.

Post reply on HN