Earlier quoted context omitted.
I just tried "analyze this audio file recording of a meeting and notes along with a transcript labeling all the speakers" (using the language from the parent's comment) and indeed Gemini 3 was significantly better than 2.5 Pro. 3 created a great "Executive Summary", identified the speakers' names, and then gave me a second by second transcript: [00:00] Greg: Hello. [00:01] X: You great? [00:02] Greg: Hi. [00:03] X: I…
Does it deduce everyone's name?
Gemini 3
241–250 of 1001 posts
Re: Gemini 3
#242Re: Gemini 3
#243Wow so the polymarket insider bet was true then.. https://old.reddit.com/r/wallstreetbets/comments/1oz6gjp/new...
These prediction markets are so ripe for abuse it's unbelievable. People need to realize there are real people on the other side of these bets. Brian Armstong, CEO of Coinbase intentionally altered the outcome of a bet by randomly stating "Bitcoin, Ethereum, blockchain, staking, Web3" at the end of an earnings call. These types of bets shouldn't be allowed.
Re: Gemini 3
#244I think I am in this AI fatigue phase. I am past all hype with models, tools and agents and back to problem and solution approach, sometimes code gen with AI , sometimes think and ask for a piece of code. But not offloading to AI and buying all the bs, waiting it to do magic with my codebase.
Re: Gemini 3
#245Re: Gemini 3
#246Re: Gemini 3
#247Antigravity uses Open-VSX and can't be configured differently even though it says it right there (setting is missing). Gemini website still only lists 2.5 Pro. Guess I will just stick to Claude.
Re: Gemini 3
#248Earlier quoted context omitted.
I like to ask "Make a pacman game in a single html page". No model has ever gotten a decent game in one shot. My attempt with Gemini3 was no better than 2.5.
Your benchmarks should not involve IP.
Re: Gemini 3
#249I absolutely LOVE that Google themselves drew a sharp distinction here.
Re: Gemini 3
#250Earlier quoted context omitted.
At this point I'm surprised they haven't been training on thousands of professionally-created SVGs of pelicans on bicycles.
i think anything that makes it clear they've done that would be a lot worse PR than failing the pelican test would ever be.
Secondly, benchmarks are public data, and these models are trained on such large amounts of it that it would be impractical to ensure that some benchmark data is not part of the training set. And even if it's not, it would be safe to assume that engineers building these models would test their performance on all kinds of benchmarks, and tweak them accordingly. This happens all the time in other industries as well.
So the pelican riding a bicycle test is interesting, but it's not a performance indicator at this point.