Live data from Hacker News

Gemini 3 Deep Think

blog.google

651–660 of 722 posts

Re: Gemini 3 Deep Think

#651
post #39

The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/

>"The pelican riding a bicycle is excellent. I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/"

Yeah this is nuts. First real step-change we've seen since Claude 3.5 in '24.

Re: Gemini 3 Deep Think

#652
post #598

Earlier quoted context omitted.

The other SVGs I tried from my private collection of prompts were all similarly impressive.

Is there a way you can showcase a few of these?

Not without people later saying "you shared that on Hacker News last year clearly the AI labs are training for it now!"

Re: Gemini 3 Deep Think

#653

Earlier quoted context omitted.

François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…

Please let’s hold M Chollet to account, at least a little. He launched ARC claiming transformer architectures could never do it and that he thought solving it would be AGI. And he was smug about it. ARC 2 had a very similar launch. Both have been crushed in far less time without significantly different architectures than he predicted. It’s a hard test! And novel, and worth continuing to iterate on. But it was not lau…

Here is what the original paper for ARC-AGI-1 said in 2019:

> Our definition, formal framework, and evaluation guidelines, which do not capture all facets of intelligence, were developed to be actionable, explanatory, and quantifiable, rather than being descriptive, exhaustive, or consensual. They are not meant to invalidate other perspectives on intelligence, rather, they are meant to serve as a useful objective function to guide research on broad AI and general AI [...]

> Importantly, ARC is still a work in progress, with known weaknesses listed in [Section III.2]. We plan on further refining the dataset in the future, both as a playground for research and as a joint benchmark for machine intelligence and human intelligence.

> The measure of the success of our message will be its ability to divert the attention of some part of the community interested in general AI, away from surpassing humans at tests of skill, towards investigating the development of human-like broad cognitive abilities, through the lens of program synthesis, Core Knowledge priors, curriculum optimization, information efficiency, and achieving extreme generalization through strong abstraction.

Re: Gemini 3 Deep Think

#654

Earlier quoted context omitted.

I’m someone who’d like to deploy a lot more workers than I want to manage. Put another way, I’m on the capital side of the conversation. The good news for labor that has experience and creativity is that it just started costing 1/100,000 what it used to to get on that side of the equation.

lmao, you are an idealistic moron. If llms can replace labor at 1/100k of the cost (lmfao) why are you looking to "deploy" more workers? So are you trying to say if I have $100.00 in tokens I have the equivalent of $10mm in labor potential.... What kind of statement is this? This is truly the dumbest statement I've ever seen on this site for too many reasons to list. You people sound like NFT people in 2021 telling p…

I upvoted your comment. Love the confidence. I’ve self funded full venture studios - so I have a pretty good take on costs of innovation. You might say I was poor at deploying innovation capital; you might be right!

Anyway 100k is hyperbolic. But I’d argue just one order of magnitude. Claude max can do many things better than my last (really great) team, and is worse at some things - creative output, relationship building and conference attending most notably. It’s also much faster at the things it is good at. Like 20-50x faster than a person or team.

If I had another venture studio I’d start with an agent first, and fill in labor in the gaps. The costs are wildly different.

Back to you though - who hurt you? Your writing makes me think you are young. You have been given literal super power force extension tech from aliens this year, why not be excited at how much more you can build?

Re: Gemini 3 Deep Think

#657
post #312

Earlier quoted context omitted.

Per BalatroBench, gemini-3-pro-preview makes it to round (not ante) 19.3 ± 6.8 on the lowest difficulty on the deck aimed at new players. Round 24 is ante 8's final round. Per BalatroBench, this includes giving the LLM a strategy guide, which first-time players do not have. Gemini isn't even emitting legal moves 100% of the time.

It beats ante eight 9 times out of 15 attempts. I do consider 60% winning chance very good for a first time player. The average is only 19.3 rounds because there is a bugged run where Gemini beats round 6 but the game bugs out when it attempts to sell Invisible Joker (a valid move)[0]. That being said, Gemini made a big mistake in round 6 that would have costed it the run at higher difficulty. [0]: given the existenc…

Are there benchmarks if we allow the LLM to practice and study the game?

Re: Gemini 3 Deep Think

#658
post #652

Earlier quoted context omitted.

Is there a way you can showcase a few of these?

Not without people later saying "you shared that on Hacker News last year clearly the AI labs are training for it now!"

Couldn't you just make up new combinations, or new caveats indefinitely to mitigate that? It would be nice to see maybe 3-4 good examples for validation. I'd do it myself, but I don't have $200 to play around with this model.

Re: Gemini 3 Deep Think

#659
post #294

Earlier quoted context omitted.

[flagged]

"Lunar New Year" is perhaps over-general, since there are non-Asian lunar calendars, such as the Hebrew and Islamic calendars. That said, "Lunar New Year" is probably as good a compromise as any, since we have other names for the Hebrew and Islamic New Years.

There's more than one Asian lunar calendar: https://news.ycombinator.com/item?id=46996396.

The Islamic calendar originated in Arabia. Calling it an Asian lunar calendar wouldn't be inaccurate.

Re: Gemini 3 Deep Think

#660

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

I mean, remember when ARC 1 was basically solved, and then ARC 2 (which is even easier for humans) came out, and all of the sudden the same models that were doing well on ARC 1 couldn’t even get 5% on ARC 2? Not convinced this isn’t data leakage.
Post reply on HN