Will run extended benchmarks later, let me know if you want to see actual data.
Gemini 3
111–120 of 1001 posts
Re: Gemini 3
#112Re: Gemini 3
#113It generated a quite cool pelican on a bike: https://imgur.com/a/yzXpEEh
2026: cure cancer
Re: Gemini 3
#114Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...
Why is this particular benchmark important?
Re: Gemini 3
#115Re: Gemini 3
#116Re: Gemini 3
#117I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…
How can you be sure that your benchmark is meaningful and well designed? Is the only thing that prevents a benchmark from being meaningful publicity?
I think my benchmark is well designed. It's well designed because it's a generalization of a problem I've consistently had with LLMs on my code. Insofar that it encapsulates my coding preferences and communication style, that's the proper benchmark for me.
Re: Gemini 3
#118Just generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed. Will run extended benchmarks later, let me know if you want to see actual data.
Re: Gemini 3
#119Re: Gemini 3
#120Not to be a negative nelly, but these numbers are definitely inflated due to Google literally pushing their AI into everything they can, much like M$. Can't even search google without getting an AI response. Surely you can't claim those numbers are legit.