Live data from Hacker News

Gemini 3 Pro Model Card [pdf]

storage.googleapis.com

311–320 of 359 posts

Re: Gemini 3 Pro Model Card [pdf]

#311

Earlier quoted context omitted.

How do they hold back questions in practice though? These are hosted models. To ask the question is to reveal it to the model team.

They pinky swear not to store and use the prompts and data lol

A legally binding pinky swear LOL

Re: Gemini 3 Pro Model Card [pdf]

#312

Earlier quoted context omitted.

Google was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.

Bard was horrible compared to the competition of the time. Gemini 1.0 was strictly worse than GPT-3.5 and was unusable due to "safety" features. Google followed that up with 1.5 which was still worse than GPT-3.5 and unbelievably far behind GPT-4. At this same time Google had their "black nazi" scandals. With Gemini 2.0 finally had a model that was at least useful for OCR and with their fash series a model that, whil…

> their fash series

Unfortunate typo.

Re: Gemini 3 Pro Model Card [pdf]

#313

There needs to be a sycophancy benchmark in these comparisons. More baseless praise and false agreement = lower score.

I care very little about model personality outside of sycophancy. The thing about gemini is that it's notorious for its low self esteem. Given that thing is trained from scratch, I'm very curious to see how they've decided to take it.

Sonnet-4.5 has the lowest self esteem of any model I've used. Gemini frequently argues with me.

Re: Gemini 3 Pro Model Card [pdf]

#314

Earlier quoted context omitted.

Does Google's team not proofread this stuff? Or maybe is this an early draft that wasn't meant to be released?

It was generated by an LLM like everything else these days.

LLMs don't make typos.

Re: Gemini 3 Pro Model Card [pdf]

#315

These model cards tell me nothing. I want to know the exact data a model was trained on. Otherwise, how can I safely use it for generating texts that I show to children? Etc.etc.

The data is everything you've ever heard of, and obviously contains things you wouldn't show to children, since that'd include NYT war journalism stories.

Re: Gemini 3 Pro Model Card [pdf]

#316

Earlier quoted context omitted.

I think Anthropic is reading the room, and just going to go hard on being "the" coding model. I suppose they feel that if they can win that, they can get an ROI without having to do full blown multimodality at the highest level. It's probably pretty liberating, because you can make a "spikey" intelligence with only one spike to really focus on.

Codex has been good enough to me and it’s much cheaper. I code non-trivial stuff with it like multi-threaded code and at least for my style of AI coding which is to do fairly small units of work with multiple revisions it is good enough for me to not to even consider the competition. Just giving you a perspective on how the benchmarks might not be important at all for some people and how Claude may have a difficult t…

>> Codex has been good enough to me and it’s much cheaper.

It may be cheaper but it's much, much slower, which is a total flow killer in my experience.

Re: Gemini 3 Pro Model Card [pdf]

#317
post #62

Interesting to see on page 2 the reference to ML pathways [1]. Looks like a multi layer mixture of experts. Is this common ? [1] https://blog.google/technology/ai/introducing-pathways-next-...

Pathways, I understand, is more so these days just the name for their training orchestrator for doing distributed JAX stuff - https://github.com/google/pathways-job

Re: Gemini 3 Pro Model Card [pdf]

#318
post #12

It says it's been trained from scratch. I wonder if it will have the same undescribable magic that makes me spend an hour every day with 2.5. I really love the results I can get with 2.5 pro. Google eventually limiting aistudio will be a sad day. Also I really hoped for a 2M+ context. I'm living on the context edge even with 1M.

buy a pixel and you get it basically unlimited for free for a year ;)

or a Chromebook is a good choice too considering price

Re: Gemini 3 Pro Model Card [pdf]

#319
post #249

Earlier quoted context omitted.

Wow. They must have had some major breakthrough. Those scores are truly insane. O_O Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Comput…

The problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.

I don't think any of these companies are that reductive and short-sighted to try to game the system. However, Goodhart's Law comes into play. I am sure they have their own metrics that arr much more detailed than these benchmarks, but the fact remains LLMs will be tuned according to elements that are deterministically measurable.

Re: Gemini 3 Pro Model Card [pdf]

#320
post #165

Earlier quoted context omitted.

Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the site refers to, adds a specific agentic harness, whereas the Gemini 3 Pro seems to be on a standard eval harness. It would be interesting to see the apples-to-apples figure, i.e. with Google's best harness alongside Codex CLI.

All evals on Terminal Bench require some harness. :) Or "Agent", as Terminal Bench calls it. Presumably the Gemini 3 are using Gemini CLI. What do you mean by "standard eval harness"?

I think the point is that it looks like Gemini 3 was only tested with the generic "Terminus 2", whereas Codex was tested with the Codex CLI.
Post reply on HN