Live data from Hacker News

Gemini 3

blog.google

191–200 of 1001 posts

Re: Gemini 3

#191

Just generated a bunch of 3D CAD models using Gemini 3.0 to see how it compares in spatial understanding and it's heaps better than anything currently out there - not only intelligence but also speed. Will run extended benchmarks later, let me know if you want to see actual data.

I'm not familiar enough with CAD what type of format is it?

When I see CAD, I always think of Casting Assistant Device.

Re: Gemini 3

#192
post #78
post #65

My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.

What prompt do you use for that?

I just tried "analyze this audio file recording of a meeting and notes along with a transcript labeling all the speakers" (using the language from the parent's comment) and indeed Gemini 3 was significantly better than 2.5 Pro.

3 created a great "Executive Summary", identified the speakers' names, and then gave me a second by second transcript:

    [00:00] Greg: Hello.
    [00:01] X: You great?
    [00:02] Greg: Hi.
    [00:03] X: I'm X.
    [00:04] Y: I'm Y.
    ...
Super impressive!

Re: Gemini 3

#193
post #126

Earlier quoted context omitted.

Your benchmarks should not involve IP.

Why? This seems like a reasonable task to benchmark on.

Sure, reasonable to benchmark on if your goal is to find out which companies are the best at stealing the hard work of some honest, human souls.

Re: Gemini 3

#194
post #65

My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.

Parakeet TDT v3 would be really good at that

Re: Gemini 3

#195
post #82

Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...

Does anyone trust benchmarks at this point? Genuine question. Isn't the scientific consensus that they are broken and poor evaluation tools?

I make my own automated benchmarks

Re: Gemini 3

#196

Earlier quoted context omitted.

I didn't tell you what you should think about the model. All I said is that you should have your own benchmark. I think my benchmark is well designed. It's well designed because it's a generalization of a problem I've consistently had with LLMs on my code. Insofar that it encapsulates my coding preferences and communication style, that's the proper benchmark for me.

I asked a semi related question in a different thread [0] -- is the basic idea behind your benchmark that you specifically keep it secret to use it as an "actually real" test that was definitely withheld from training new LLMs? I've been thinking about making/publishing a new eval - if it's not public, presumably LLMs would never get better at them. But is your fear that generally speaking, LLMs tend to (I don't want…

> if it's not public, presumably LLMs would never get better at them.

Why? This is not obvious to me at all.

Re: Gemini 3

#197
The first paragraph is pure delusion. Why do investors like delusional CEOs so much? I would take it as a major red flag.

Re: Gemini 3

#198

Earlier quoted context omitted.

I'm primarily reacting to the other threads, like the one that leaked the system card early. And, perhaps unfairly, Twitter as well.

You might not believe this, but there are a lot of people (me included) that were extremely excited about the Gemini 3 release and are pleased to see the SOTA benchmark results, and this is reflected in the comments.

I definitely believe it--I'm not a total AI hater. The jump on the screen usage benchmark is really exciting in that it might substantially help computer-use agentic workflows.

That said, I think there is too much a pattern with recent model releases around what appears to me to be astroturfing to get to HN front page. Of course that doesn't preclude many organic comments that are excited too!

A bit of both always happens. But given how important these model releases are to justify the capex and levels of investment, I think it is pretty clear the various "front pages" of our internet are manipulated. The incentive is just too strong not to.

Re: Gemini 3

#199

Grok got to hold the top spot of LMArena-text for all of ~24 hours, good for them [1]. With stylecontrol enabled, that is. Without stylecontrol, gemini held the fort. [1] https://lmarena.ai/leaderboard/text

Is it just me or is that link broken because of the cloudflare outage?

Edit: nvm it looks to be up for me again

Re: Gemini 3

#200
post #87

I've been so happy to see Google wake up. Many can point to a long history of killed products and soured opinions but you can't deny theyve been the great balancing force (often for good) in the industry. - Gmail vs Outlook - Drive vs Word - Android vs iOS - Worklife balance and high pay vs the low salary grind of before. Theyve done heaps for the industry. Im glad to see signs of life. Particularly in their P/E whic…

Forgot to mention absolutely milking every ounce of their users attention with Youtube, plus forcing Shorts!
Post reply on HN