Live data from Hacker News

Gemini 3 Deep Think

blog.google

111–120 of 722 posts

Re: Gemini 3 Deep Think

#111
post #88

Does anyone actually use Gemini 3 now? I cant stand its sleek salesy way of introduction, and it doesnt hold to instructions hard – makes it unapplicable for MECE breakdowns or for writing.

I do. It's excellent when paired with an MCP like context7.

Re: Gemini 3 Deep Think

#112

Earlier quoted context omitted.

For every combination of animal and vehicle? Very unlikely. The beauty of this benchmark is that it takes all of two seconds to come up with your own unique one. A seahorse on a unicycle. A platypus flying a glider. A man’o’war piloting a Portuguese man of war. Whatever you want.

No, not every combination. The question is about the specific combination of a pelican on a bicycle. It might be easy to come up with another test, but we're looking at the results from a particular one here.

More likely you would just train for emitting svg for some description of a scene and create training data from raster images.

Re: Gemini 3 Deep Think

#113
post #15

Google is absolutely running away with it. The greatest trick they ever pulled was letting people think they were behind.

But wait two hours for what OpenAI has! I love the competition and how someone just a few days ago was telling how ARC-AGI-2 was proof that LLMs can't reason. The goalposts will shift again. I feel like most of human endeavor will soon be just about trying to continuously show that AI's don't have AGI.

Re: Gemini 3 Deep Think

#114

Gemini was awesome and now it’s garbage. It’s impossible for it to do anything but cut code down, drop features, lose stuff and give you less than the code you put in. It’s puzzling because it spent months at the head of the pack now I don’t use it at all because why do I want any of those things when I’m doing development. I’m a paid subscriber but there’s no point any more I’ll spend the money on Claude 4.6 instead…

I never found it useful for code. It produced garbage littered with gigantic comments.

Me: Remove comments

Literally Gemini: // Comments were removed

Re: Gemini 3 Deep Think

#115

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

I'm excited for the big jump in ARC-AGI scores from recent models, but no one should think for a second this is some leap in "general intelligence".

I joke to myself that the G in ARC-AGI is "graphical". I think what's held back models on ARC-AGI is their terrible spatial reasoning, and I'm guessing that's what the recent models have cracked.

Looking forward to ARC-AGI 3, which focuses on trial and error and exploring a set of constraints via games.

Re: Gemini 3 Deep Think

#116
post #69
post #40

Earlier quoted context omitted.

It's worth noting that you mean excellent in terms of prior AI output. I'm pretty sure this wouldn't be considered excellent from a "human made art" perspective. In other words, it's still got a ways to go! Edit: someone needs to explain why this comment is getting downvoted, because I don't understand. Did someone's ego get hurt, or what?

It depends, if you meant from a human coding an SVG "manually" the same way, I'd still say this is excellent (minus the reflection issue). If you meant a human using a proper vector editor, then yeah.

maybe you're a pro vector artist but I couldn't create such a cool one myself in illustrator tbh

Re: Gemini 3 Deep Think

#119
post #70

I can't shake of the feeling that Googles Deep Think Models are not really different models but just the old ones being run with higher number of parallel subagents, something you can do by yourself with their base model and opencode.

And after i do that, how do i combine the output of 1000 subagents into one output? (Im not being snarky here, i think it's a nontrivial problem)

Start with 1024 and use half the number of agents each turn to distill the final result.

Re: Gemini 3 Deep Think

#120

Is it me or is the rate of model release is accelerating to an absurd degree? Today we have Gemini 3 Deep Think and GPT 5.3 Codex Spark. Yesterday we had GLM5 and MiniMax M2.5. Five days before that we had Opus 4.6 and GPT 5.3. Then maybe two weeks I think before that we had Kimi K2.5.

Anthropic took the day off to do a $30B raise at a $380B valuation.
Post reply on HN