Live data from Hacker News

Gemini 3

blog.google

241–250 of 1001 posts

Re: Gemini 3

#241

Earlier quoted context omitted.

I just tried "analyze this audio file recording of a meeting and notes along with a transcript labeling all the speakers" (using the language from the parent's comment) and indeed Gemini 3 was significantly better than 2.5 Pro. 3 created a great "Executive Summary", identified the speakers' names, and then gave me a second by second transcript: [00:00] Greg: Hello. [00:01] X: You great? [00:02] Greg: Hi. [00:03] X: I…

Does it deduce everyone's name?

It does! I redacted them, but yes. This was a 3-person call.

Re: Gemini 3

#243
post #134

Wow so the polymarket insider bet was true then.. https://old.reddit.com/r/wallstreetbets/comments/1oz6gjp/new...

These prediction markets are so ripe for abuse it's unbelievable. People need to realize there are real people on the other side of these bets. Brian Armstong, CEO of Coinbase intentionally altered the outcome of a bet by randomly stating "Bitcoin, Ethereum, blockchain, staking, Web3" at the end of an earnings call. These types of bets shouldn't be allowed.

The point of prediction markets isn't to be fair. They are not the stock market. The point of prediction markets is to predict. They provide a monetary incentive for people who are good at predicting stuff. Whether that's due to luck, analysis, insider knowledge, or the ability to influence the result is irrelevant. If you don't want to participate in an unfair market, don't participate in prediction markets.

Re: Gemini 3

#244

I think I am in this AI fatigue phase. I am past all hype with models, tools and agents and back to problem and solution approach, sometimes code gen with AI , sometimes think and ask for a piece of code. But not offloading to AI and buying all the bs, waiting it to do magic with my codebase.

it's not AI fatigue, its that you just need to shift mode to not pay attention too much to the latest and greatest as they all leap frog each other each month. Just stick to one and ride it thru ups and downs.

Re: Gemini 3

#245
This is a really impressive release. It's probably the biggest lead we've seen from a model since the release of GPT-4. Seems likely that OpenAI rushed out GPT-5.1 to beat the Gemini 3 release, knowing that their model would underperform it.

Re: Gemini 3

#247
Gemini CLI crashes due to this bug: https://github.com/google-gemini/gemini-cli/issues/13050 and when applying the fix in the settings file I can't login with my Google account due to "The authentication did not complete successfully. The following products are not yet authorized to access your account" with useless links to completely different products (Code Assist).

Antigravity uses Open-VSX and can't be configured differently even though it says it right there (setting is missing). Gemini website still only lists 2.5 Pro. Guess I will just stick to Claude.

Re: Gemini 3

#248
post #126

Earlier quoted context omitted.

I like to ask "Make a pacman game in a single html page". No model has ever gotten a decent game in one shot. My attempt with Gemini3 was no better than 2.5.

Your benchmarks should not involve IP.

The only intellectual property here would be trademark. No copyright, no patent, no trade secret. Unless someone wants to market the test results as a genuine Pac-Man-branded product, or otherwise dilute that brand, there's nothing should-y about it.

Re: Gemini 3

#249
> Whether you’re an experienced developer or a vibe coder

I absolutely LOVE that Google themselves drew a sharp distinction here.

Re: Gemini 3

#250

Earlier quoted context omitted.

At this point I'm surprised they haven't been training on thousands of professionally-created SVGs of pelicans on bicycles.

i think anything that makes it clear they've done that would be a lot worse PR than failing the pelican test would ever be.

It would be next to impossible for anyone without insider knowledge to prove that to be the case.

Secondly, benchmarks are public data, and these models are trained on such large amounts of it that it would be impractical to ensure that some benchmark data is not part of the training set. And even if it's not, it would be safe to assume that engineers building these models would test their performance on all kinds of benchmarks, and tweak them accordingly. This happens all the time in other industries as well.

So the pelican riding a bicycle test is interesting, but it's not a performance indicator at this point.

Post reply on HN