Live data from Hacker News

Gemini 3

blog.google

111–120 of 1001 posts

Re: Gemini 3

#111
Just generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed.

Will run extended benchmarks later, let me know if you want to see actual data.

Re: Gemini 3

#112
I think I am in this AI fatigue phase. I am past all hype with models, tools and agents and back to problem and solution approach, sometimes code gen with AI , sometimes think and ask for a piece of code. But not offloading to AI and buying all the bs, waiting it to do magic with my codebase.

Re: Gemini 3

#114
post #82

Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...

Why is this particular benchmark important?

Thus far, this is one of the best objective evaluations of real world software engineering...

Re: Gemini 3

#115

Earlier quoted context omitted.

Still cheaper than Sonnet 4.5: $3/M for input and $15/M for output.

It is so impressive that Anthropic has been able to maintain this pricing still.

Because every time I try to move away I realize there’s nothing equivalent to move to.

Re: Gemini 3

#117
post #95

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

How can you be sure that your benchmark is meaningful and well designed? Is the only thing that prevents a benchmark from being meaningful publicity?

I didn't tell you what you should think about the model. All I said is that you should have your own benchmark.

I think my benchmark is well designed. It's well designed because it's a generalization of a problem I've consistently had with LLMs on my code. Insofar that it encapsulates my coding preferences and communication style, that's the proper benchmark for me.

Re: Gemini 3

#118

Just generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed. Will run extended benchmarks later, let me know if you want to see actual data.

I'm not familiar enough with CAD what type of format is it?

Re: Gemini 3

#119

Earlier quoted context omitted.

Wow, you weren't wrong...

It's the only comment referencing AGI. Seems wrong to me.

I'm primarily reacting to the other threads, like the one that leaked the system card early. And, perhaps unfairly, Twitter as well.

Re: Gemini 3

#120
> The Gemini app surpasses 650 million users per month, more than 70% of our Cloud customers use our AI, 13 million developers have built with our generative models, and that is just a snippet of the impact we’re seeing

Not to be a negative nelly, but these numbers are definitely inflated due to Google literally pushing their AI into everything they can, much like M$. Can't even search google without getting an AI response. Surely you can't claim those numbers are legit.

Post reply on HN