The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
Yeah this is nuts. First real step-change we've seen since Claude 3.5 in '24.
651–660 of 722 posts
The pelican riding a bicycle is excellent . I think it's the best I've seen. https://simonwillison.net/2026/Feb/12/gemini-3-deep-think/
Yeah this is nuts. First real step-change we've seen since Claude 3.5 in '24.
Earlier quoted context omitted.
The other SVGs I tried from my private collection of prompts were all similarly impressive.
Is there a way you can showcase a few of these?
Earlier quoted context omitted.
François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…
Please let’s hold M Chollet to account, at least a little. He launched ARC claiming transformer architectures could never do it and that he thought solving it would be AGI. And he was smug about it. ARC 2 had a very similar launch. Both have been crushed in far less time without significantly different architectures than he predicted. It’s a hard test! And novel, and worth continuing to iterate on. But it was not lau…
> Our definition, formal framework, and evaluation guidelines, which do not capture all facets of intelligence, were developed to be actionable, explanatory, and quantifiable, rather than being descriptive, exhaustive, or consensual. They are not meant to invalidate other perspectives on intelligence, rather, they are meant to serve as a useful objective function to guide research on broad AI and general AI [...]
> Importantly, ARC is still a work in progress, with known weaknesses listed in [Section III.2]. We plan on further refining the dataset in the future, both as a playground for research and as a joint benchmark for machine intelligence and human intelligence.
> The measure of the success of our message will be its ability to divert the attention of some part of the community interested in general AI, away from surpassing humans at tests of skill, towards investigating the development of human-like broad cognitive abilities, through the lens of program synthesis, Core Knowledge priors, curriculum optimization, information efficiency, and achieving extreme generalization through strong abstraction.
Earlier quoted context omitted.
I’m someone who’d like to deploy a lot more workers than I want to manage. Put another way, I’m on the capital side of the conversation. The good news for labor that has experience and creativity is that it just started costing 1/100,000 what it used to to get on that side of the equation.
lmao, you are an idealistic moron. If llms can replace labor at 1/100k of the cost (lmfao) why are you looking to "deploy" more workers? So are you trying to say if I have $100.00 in tokens I have the equivalent of $10mm in labor potential.... What kind of statement is this? This is truly the dumbest statement I've ever seen on this site for too many reasons to list. You people sound like NFT people in 2021 telling p…
Anyway 100k is hyperbolic. But I’d argue just one order of magnitude. Claude max can do many things better than my last (really great) team, and is worse at some things - creative output, relationship building and conference attending most notably. It’s also much faster at the things it is good at. Like 20-50x faster than a person or team.
If I had another venture studio I’d start with an agent first, and fill in labor in the gaps. The costs are wildly different.
Back to you though - who hurt you? Your writing makes me think you are young. You have been given literal super power force extension tech from aliens this year, why not be excited at how much more you can build?
I've been wondering for a while now: What would be the results if we had multiple LLMs run the same query and then use statistical analysis?
Earlier quoted context omitted.
Per BalatroBench, gemini-3-pro-preview makes it to round (not ante) 19.3 ± 6.8 on the lowest difficulty on the deck aimed at new players. Round 24 is ante 8's final round. Per BalatroBench, this includes giving the LLM a strategy guide, which first-time players do not have. Gemini isn't even emitting legal moves 100% of the time.
It beats ante eight 9 times out of 15 attempts. I do consider 60% winning chance very good for a first time player. The average is only 19.3 rounds because there is a bugged run where Gemini beats round 6 but the game bugs out when it attempts to sell Invisible Joker (a valid move)[0]. That being said, Gemini made a big mistake in round 6 that would have costed it the run at higher difficulty. [0]: given the existenc…
Earlier quoted context omitted.
Is there a way you can showcase a few of these?
Not without people later saying "you shared that on Hacker News last year clearly the AI labs are training for it now!"
Earlier quoted context omitted.
[flagged]
"Lunar New Year" is perhaps over-general, since there are non-Asian lunar calendars, such as the Hebrew and Islamic calendars. That said, "Lunar New Year" is probably as good a compromise as any, since we have other names for the Hebrew and Islamic New Years.
The Islamic calendar originated in Arabia. Calling it an Asian lunar calendar wouldn't be inaccurate.
Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...