Just generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed. Will run extended benchmarks later, let me know if you want to see actual data.
I'm not familiar enough with CAD what type of format is it?
Gemini 3
191–200 of 1001 posts
Re: Gemini 3
#192My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.
What prompt do you use for that?
3 created a great "Executive Summary", identified the speakers' names, and then gave me a second by second transcript:
[00:00] Greg: Hello.
[00:01] X: You great?
[00:02] Greg: Hi.
[00:03] X: I'm X.
[00:04] Y: I'm Y.
...
Super impressive!Re: Gemini 3
#193Re: Gemini 3
#194My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.
Re: Gemini 3
#195Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...
Does anyone trust benchmarks at this point? Genuine question. Isn't the scientific consensus that they are broken and poor evaluation tools?
Re: Gemini 3
#196Earlier quoted context omitted.
I didn't tell you what you should think about the model. All I said is that you should have your own benchmark. I think my benchmark is well designed. It's well designed because it's a generalization of a problem I've consistently had with LLMs on my code. Insofar that it encapsulates my coding preferences and communication style, that's the proper benchmark for me.
I asked a semi related question in a different thread [0] -- is the basic idea behind your benchmark that you specifically keep it secret to use it as an "actually real" test that was definitely withheld from training new LLMs? I've been thinking about making/publishing a new eval - if it's not public, presumably LLMs would never get better at them. But is your fear that generally speaking, LLMs tend to (I don't want…
Why? This is not obvious to me at all.
Re: Gemini 3
#197Re: Gemini 3
#198Earlier quoted context omitted.
I'm primarily reacting to the other threads, like the one that leaked the system card early. And, perhaps unfairly, Twitter as well.
You might not believe this, but there are a lot of people (me included) that were extremely excited about the Gemini 3 release and are pleased to see the SOTA benchmark results, and this is reflected in the comments.
That said, I think there is too much a pattern with recent model releases around what appears to me to be astroturfing to get to HN front page. Of course that doesn't preclude many organic comments that are excited too!
A bit of both always happens. But given how important these model releases are to justify the capex and levels of investment, I think it is pretty clear the various "front pages" of our internet are manipulated. The incentive is just too strong not to.
Re: Gemini 3
#199Grok got to hold the top spot of LMArena-text for all of ~24 hours, good for them [1]. With stylecontrol enabled, that is. Without stylecontrol, gemini held the fort. [1] https://lmarena.ai/leaderboard/text
Edit: nvm it looks to be up for me again
Re: Gemini 3
#200I've been so happy to see Google wake up. Many can point to a long history of killed products and soured opinions but you can't deny theyve been the great balancing force (often for good) in the industry. - Gmail vs Outlook - Drive vs Word - Android vs iOS - Worklife balance and high pay vs the low salary grind of before. Theyve done heaps for the industry. Im glad to see signs of life. Particularly in their P/E whic…