Live data from Hacker News

Gemini 3 Deep Think

blog.google

181–190 of 722 posts

Re: Gemini 3 Deep Think

#181

Earlier quoted context omitted.

It's not actually, look up some photos of the sun setting over the ocean. Here's an example: https://stockcake.com/i/sunset-over-ocean_1317824_81961

That’s only if the sun is above the horizon entirely.

No, it's not.

https://stockcake.com/i/serene-ocean-sunset_1152191_440307

Re: Gemini 3 Deep Think

#182

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

> . I don't think there are many people who posted their Balatro playthroughs in text form online

There are *tons* of balatro content on YouTube though, and it makes absolutely zero doubt that Google is using YouTube content to train their model.

Re: Gemini 3 Deep Think

#183
post #148

Earlier quoted context omitted.

I look forward to them trying. I'll know when the pelican riding a bicycle is good but the ocelot riding a skateboard sucks.

But they could just train on an assortment of animals and vehicles. It's the kind of relatively narrow domain where NNs could reasonably interpolate.

The idea that an AI lab would pay a small army of human artists to create training data for $animal on $transport just to cheat on my stupid benchmark delights me.

Re: Gemini 3 Deep Think

#184

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

But... there's Deepseek v3.2 in your link (rank 7)

Re: Gemini 3 Deep Think

#185
post #151

Earlier quoted context omitted.

Gemini's UX (and of course privacy cred as with anything Google) is the worst of all the AI apps. In the eyes of the Common Man, it's UI that will win out, and ChatGPT's is still the best.

Google privacy cred is ... excellent? The worst data breach I know of them having was a flaw that allowed access to names and emails of 500k users.

If you consider "privacy" to be 'a giant corporation tracks every bit of possible information about you and everyone else'?

Re: Gemini 3 Deep Think

#186
The problem here is that it looks like this is released with almost no real access. How are people using this without submitting to a $250/mo subscription?

Re: Gemini 3 Deep Think

#187

Earlier quoted context omitted.

It's a useless meaningless benchmark though, it just got a catchy name, as in, if the models solve this it means they have "AGI", which is clearly rubbish. Arc-AGI score isn't correlated with anything useful.

how would we actually objectively measure a model to see if it is AGI if not with benchmarks like arc-AGI?

Give it a prompt like

>can u make the progm for helps that with what in need for shpping good cheap products that will display them on screen and have me let the best one to get so that i can quickly hav it at home

And get back an automatic coupon code app like the user actually wanted.

Re: Gemini 3 Deep Think

#188

Earlier quoted context omitted.

Even before this, Gemini 3 has always felt unbelievably 'general' for me. It can beat Balatro (ante 8) with text description of the game alone[0]. Yeah, it's not an extremely difficult goal for humans, but considering: 1. It's an LLM, not something trained to play Balatro specifically 2. Most (probably >99.9%) players can't do that at the first attempt 3. I don't think there are many people who posted their Balatro p…

> . I don't think there are many people who posted their Balatro playthroughs in text form online There are * tons * of balatro content on YouTube though, and it makes absolutely zero doubt that Google is using YouTube content to train their model.

Yeah, or just the steam text guides would be a huge advantage.

I really doubt it's playing completely blind

Re: Gemini 3 Deep Think

#189
post #53

Earlier quoted context omitted.

https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.

Why 5-10 years? At current rates, price per equivalent output is dropping at 99.9% over 5 years. That's basically $0.01 in 5 years. Does it really need to be that cheap to be worth it? Keep in mind, $0.01 in 5 years is worth less than $0.01 today.

Wow that's incredible! Could you show your work?

Re: Gemini 3 Deep Think

#190
post #183

Earlier quoted context omitted.

But they could just train on an assortment of animals and vehicles. It's the kind of relatively narrow domain where NNs could reasonably interpolate.

The idea that an AI lab would pay a small army of human artists to create training data for $animal on $transport just to cheat on my stupid benchmark delights me.

When you're spending trillions on capex, paying a couple of people to make some doodles in SVGs would not be a big expense.
Post reply on HN