Live data from Hacker News

Gemini 3

blog.google

551–560 of 1001 posts

Re: Gemini 3

#551
post #487

Here are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/

I was interested (and slightly disappointed) to read that the knowledge cutoff for Gemini 3 is the same as for Gemini 2.5: January 2025. I wonder why they didn't train it on more recent data. Is it possible they use the same base pre-trained model and just fine-tuned and RL-ed it better (which, of course, is where all the secret sauce training magic is these days anyhow)? That would be odd, especially for a major ver…

The model card says: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...

> This model is not a modification or a fine-tune of a prior model.

I'm curious why they decided not to update the training data cutoff date too.

Re: Gemini 3

#552
post #506

I asked Gemini to write "a comment response to this thread. I want to start an intense discussion". Gemini 3: The cognitive dissonance in this thread is staggering. We are sitting here cheering for a model that effectively closes the loop on Google’s total information dominance, while simultaneously training our own replacements. Two things in this thread should be terrifying, yet are being glossed over in favor of "…

Gotta hand it to gemini, those are some top notch points

The "Model card leak" point is worth negative points though, as it's clearly a misreading of reality.

Re: Gemini 3

#553
post #501

Earlier quoted context omitted.

Imho Gemini 2.5 was by far the better model on non-trivial tasks.

To this day, I still don't understand why Claude gets more acclaim for coding. Gemini 2.5 consistently outperformed Claude and ChatGPT mostly because of the much larger context.

Different styles of usage? I see Gemini praised for being able to feed the whole project and ask changes. Which is cool and all but... I never do that. Claude for me is better for specific modifications to specific parts of the app. There's a lot of context behind what's "better".

Re: Gemini 3

#554
post #492
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

Just to clarify the context for future readers: the latest problem at the moment is #970: https://projecteuler.net/problem=970

[deleted]

Re: Gemini 3

#555

Earlier quoted context omitted.

I like to ask "Make a pacman game in a single html page". No model has ever gotten a decent game in one shot. My attempt with Gemini3 was no better than 2.5.

Something else to consider. I often have much better success with something like: Create a prompt that creates a specification for a pacman game in a single html page. Consider edge cases and key implementation details that result in bugs. , execute prompt. It will often yield a much better result than one generic prompt. Now that models are trained on how to generate prompts for themselves this is quite productive.…

I thought this kind of chaining was already part of these systems.

Re: Gemini 3

#556

Earlier quoted context omitted.

>>benchmarks are meaningless No they’re not. Maybe you mean to say they don’t tell the whole story or have their limitations, which has always been the case. >>my fairly basic python benchmark I suspect your definition of “basic” may not be consensus. Gpt-5 thinking is a strong model for basic coding and it’d be interesting to see a simple python task it reliably fails at.

they are not meaningless, but when you work a lot with LLMs and know them VERY well, then a few varied, complex prompts tell you all you need to know about things like EQ, sycophancy, and creative writing. I like to compare them using chathub using the same prompts Gemini still calls me "the architect" in half of the prompts. It's very cringe.

I get that one can perhaps have an intuition about these things, but doesn't this seem like a somewhat flawed attitude to have all things considered? That is, saying something to the effect of "well I know its not too sycophantic, no measurement needed, I have some special prompts of my own and it passed with flying colors!" just sounds a little suspect on first pass, even if its not like totally unbelievable I guess.

Re: Gemini 3

#557

Earlier quoted context omitted.

This is an understandable, but simplistic way of looking at the world. Are you also gonna blame Apple for mining for rare earths, because they made a successful product that requires exotic materials which needs to be mined from earth? How about hundreds of thousands of factory workers that are being subjected to inhumane conditions to assemble iPhones each year? For every "OMG, internet is filled with ads", people a…

Yes, we're absolutely holding Apple accountable for outsourcing jobs, degrading the US markets, using slave and child labor, laundering cobalt from illegal "artisanal" mines in the DRC, and whitewashing what they do by using corporate layering and shady deals to put themselves at sufficient degrees of separation from problematic labor and sources to do good PR, but not actually decoupling at all. I also hold American…

> laundering cobalt from illegal "artisanal" mines in the DRC

They don't, all cobalt in Apple products is recycled.

> and whitewashing what they do by using corporate layering and shady deals to put themselves at sufficient degrees of separation from problematic labor and sources to do good PR, but not actually decoupling at all.

They don't, Apple audits their entire supply chain so it wouldn't hide anything if something moved to another subcontractor.

Re: Gemini 3

#558

Earlier quoted context omitted.

They've poisoned the internet with their monopoly on advertising, the air pollution of the online world, which is an transgression that far outweighs any good they might have done. Much of the negative social effects of being online come from the need to drive more screen time, more engagement, more clicks, and more ad impressions firehosed into the faces of users for sweet, sweet, advertiser money. When Google final…

This is an understandable, but simplistic way of looking at the world. Are you also gonna blame Apple for mining for rare earths, because they made a successful product that requires exotic materials which needs to be mined from earth? How about hundreds of thousands of factory workers that are being subjected to inhumane conditions to assemble iPhones each year? For every "OMG, internet is filled with ads", people a…

> How about hundreds of thousands of factory workers that are being subjected to inhumane conditions to assemble iPhones each year?

That would be bad if it happened, which is why it doesn't happen. Working in a factory isn't an inhumane condition.

Re: Gemini 3

#559

I asked Gemini to write "a comment response to this thread. I want to start an intense discussion". Gemini 3: The cognitive dissonance in this thread is staggering. We are sitting here cheering for a model that effectively closes the loop on Google’s total information dominance, while simultaneously training our own replacements. Two things in this thread should be terrifying, yet are being glossed over in favor of "…

> We are cheering for a product sold back to us at a 60% markup (input costs up to $2.00/M) that was built on our own private correspondence.

That feels like something between a hallucination and an intentional fallacy that popped up because you specifically said "intense discussion". The increase is 60% on input tokens from the old model, but it's not a markup, and especially not "sold back to us at X markup".

I've seen more and more of these kinds of hallucinations as these models seem to be RL'd to not be a sycophant, they're slowly inching into the opposite direction where they tell small fibs or embellish in a way that seems like it's meant to add more weight to their answers.

I wonder if it's a form of reward hacking, since it trades being maximally accurate for being confident, and that might result in better rewards than being accurate and precise

Re: Gemini 3

#560
post #487

Here are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/

It's interesting that you mentioned on a recent post that saturation on the pelican benchmark isn't a problem because it's easy to test for generalization. But now looking at your updated benchmark results, I'm not sure I agree. Have the main labs been climbing the Pelican on a bike hill in secret this whole time?
Post reply on HN