Live data from Hacker News

Gemini 3 Deep Think

blog.google

191–200 of 722 posts

Re: Gemini 3 Deep Think

#191
post #84

Less than a year to destroy Arc-AGI-2 - wow.

But why only a +0.5% increase for MMMU-Pro?

Its possibly label noise. But you can't tell from a single number.

You would need to check to see if everyone is having mistakes on the same 20% or different 20%. If its the same 20% either those questions are really hard, or they are keyed incorrectly, or they aren't stated with enough context to actually solve the problem.

It happens. Old MMLU non pro had a lot of wrong answers. Simple things like MNIST have digits labeled incorrect or drawn so badly its not even a digit anymore.

Re: Gemini 3 Deep Think

#192
post #68

Earlier quoted context omitted.

Isn’t the point of ARC that you can’t train against it? Or doesn’t it achieve that goal anymore somehow?

How can you make sure of that? AFAIK, these SOTA models run exclusively on their developers hardware. So any test, any benchmark, anything you do, does leak per definition. Considering the nature of us humans and the typical prisoners dilemma, I don't see how they wouldn't focus on improving benchmarks even when it gets a bit... shady? I tell this as a person who really enjoys AI by the way.

[deleted]

Re: Gemini 3 Deep Think

#193
post #171
post #53

Earlier quoted context omitted.

https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.

What’s reasonable? It’s less than minimum hourly wage in some countries.

Burned in seconds.

Re: Gemini 3 Deep Think

#194

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

[dead]

Re: Gemini 3 Deep Think

#195
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

Their models might be impressive, but their products absolutely suck donkey balls. I’ve given Gemini web/cli two months and ran away back to ChatGPT. Seriously, it would just COMPLETELY forget context mid dialog. When asked about improving air quality it just gave me a list of (mediocre) air purifiers without asking for any context whatsoever, and I can list thousands of conversations like that. Shopping or comparing…

I agree. On top of that, in true Google style, basic things just don't work.

Any time I upload an attachment, it just fails with something vague like "couldn't process file". Whether that's a simple .MD or .txt with less than 100 lines or a PDF. I tried making a gem today. It just wouldn't let me save it, with some vague error too.

I also tried having it read and write stuff to "my stuff" and Google drive. But it would consistently write but not be able to read from it again. Or would read one file from Google drive and ignore everything else.

Their models are seriously impressive. But as usual Google sucks at making them work well in real products.

Re: Gemini 3 Deep Think

#197

Earlier quoted context omitted.

Anthropic took the day off to do a $30B raise at a $380B valuation.

Most ridiculous valuation in the history of markets. Cant wait to watch these compsnies crash snd burn when people give up on the slot machine.

Why? They had $10+ billion arr run rate in 2025 trippeled from 2024 I mean 30x is a lot but also not insane at that growth rate right?

Re: Gemini 3 Deep Think

#198
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

But wait two hours for what OpenAI has! I love the competition and how someone just a few days ago was telling how ARC-AGI-2 was proof that LLMs can't reason. The goalposts will shift again. I feel like most of human endeavor will soon be just about trying to continuously show that AI's don't have AGI.

> I feel like most of human endeavor will soon be just about trying to continuously show that AI's don't have AGI.

I think you overestimate how much your average person-on-the-street cares about LLM benchmarks. They already treat ChatGPT or whichever as generally intelligent (including to their own detriment), are frustrated about their social media feeds filling up with slop and, maybe, if they're white-collar, worry about their jobs disappearing due to AI. Apart from a tiny minority in some specific field, people already know themselves to be less intelligent along any measurable axis than someone somewhere.

Re: Gemini 3 Deep Think

#199

Do we get any model architecture details like parameter size etc.? Few months back, we used to talk more on this, now it's mostly about model capabilities.

I'm honestly not sure what you mean? The frontier labs have kept arch as secrets since gpt3.5

At the very least gemini 3's flyer claims 1T parameters.

Re: Gemini 3 Deep Think

#200
post #115

Earlier quoted context omitted.

I'm excited for the big jump in ARC-AGI scores from recent models, but no one should think for a second this is some leap in "general intelligence". I joke to myself that the G in ARC-AGI is "graphical". I think what's held back models on ARC-AGI is their terrible spatial reasoning, and I'm guessing that's what the recent models have cracked. Looking forward to ARC-AGI 3, which focuses on trial and error and explorin…

The average ARC AGI 2 score for a single human is around 60%. "100% of tasks have been solved by at least 2 humans (many by more) in under 2 attempts. The average test-taker score was 60%." https://arcprize.org/arc-agi/2/

Worth keeping in mind that in this case the test takers were random members of the general public. The score of e.g. people with bachelor's degrees in science and engineering would be significantly higher.
Post reply on HN