Gemini 3 Deep Think
511–520 of 722 posts
Re: Gemini 3 Deep Think
#512Here is the methodologies for all the benchmarks: https://storage.googleapis.com/deepmind-media/gemini/gemini_... The arc-agi-2 score (84.6%) is from the semi-private eval set. If gemini-3-deepthink gets above 85% on the private eval set, it will be considered "solved" >Submit a solution which scores 85% on the ARC-AGI-2 private evaluation set and win $700K. https://arcprize.org/guide#overview
Re: Gemini 3 Deep Think
#513Earlier quoted context omitted.
Aren't we saying "lunar new year" now?
I don't think so; there are different lunar calendars.
Re: Gemini 3 Deep Think
#514Earlier quoted context omitted.
Ok, here I am living in the real world finding these models have advanced incredibly over the past year for coding. Benchmaxxing exists, but that’s not the only data point. It’s pretty clear that models are improving quickly in many domains in real world usage.
I agree completely. I think we're in alignment with Elon Musk who says that AI will bypass coding entirely and create the binary directly. It's going to be an exciting year.
Re: Gemini 3 Deep Think
#515Earlier quoted context omitted.
When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.
Vetting them for the potential for whistleblowing might be a bit more involved. But conspiracy theories have an advantage because the lack of evidence is evidence for the theory.
This would just be one more checkbox buried in hundreds of pages of requests, and compared to plenty of other ethical grey areas like copyright laundering with actual legal implications, leaking that someone was asked to create a few dozen pelican images seems like it would be at the very bottom of the list of reputational risks.
Re: Gemini 3 Deep Think
#516Earlier quoted context omitted.
> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…
Isn’t that super intelligence not AGI? Feels like these benchmarks continue to move the goalposts.
AGI without superintelligence is quite difficult to adjudicate because any time it fails at an "easy" task there will be contention about the criteria.
Re: Gemini 3 Deep Think
#517Re: Gemini 3 Deep Think
#518Earlier quoted context omitted.
Strange, because I could not for the life of me get Gemini 3 to follow my instructions the other day to work through an example with a table, Claude got it first try.
Claude is king for agentic workflows right now because it’s amazing at tool calling and following instructions well (among other things)
Re: Gemini 3 Deep Think
#519Earlier quoted context omitted.
> which is why only "ARC-AGI Certified" results using a secret problem set really matter. The 84.6% is certified and that's a pretty big deal. So, I'd agree if this was on the true fully private set, but Google themselves says they test on only the semi-private: > ARC-AGI-2 results are sourced from the ARC Prize website and are ARC Prize Verified. The set reported is v2, semi-private ( https://storage.googleapis.com/…
Chollet himself says "We certified these scores in the past few days." https://x.com/fchollet/status/2021983310541729894 . The ARC-AGI papers claim to show that training on a public or semi-private set of ARC-AGI problems to be of very limited value in passing a private set. none of ARC-AGI can possibly be valid. So, before "public, semi-private or private" answers leaking or 'benchmaxing' on them can even matter - y…
But I think such quibbling largely misses the point. The goal is really just to guarantee that the test isn't unintentionally trained on. For that, semi-private is sufficient.
Re: Gemini 3 Deep Think
#520Earlier quoted context omitted.
I think it is because of the Chinese new year. The Chinese labs like to publish their models arround the Chinese new year, and the US labs do not want to let a DeepSeek R1 (20 January 2025) impact event happen again, so i guess they publish models that are more capable then what they imagine Chinese labs are yet capable of producing.
I guess. Deepseek v3 was released on boxing day a month prior https://api-docs.deepseek.com/news/news1226
It was R1 with its RL-training that made the news and crashed the srock market.