Live data from Hacker News

Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

blog.jcz.dev

21–30 of 97 posts

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#22
post #4

"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?

It might get things wrong on purpose, but deep down it knows what it's doing

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#23
post #12

1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?

“Gemini 3 Pro was often overloaded, which produced long spans of downtime that 2.5 Pro experienced much less often”

I was unclear if this meant that the API was overloaded or if he was on a subscription plan and had hit his limit for the moment. Although I think that the Gemini plans just use weekly limits, so I guess it must be API.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#24
post #14

Earlier quoted context omitted.

My issue with this is that the LLM could just be roleplaying that it doesn't know.

To test would just need to edit the rom and switch around the solution. Not sure how complicated that is, likely depends on the rom system.

I don't know why people still get wrapped around the axle of "training data".

Basically every benchmark worth it's salt uses bespoke problems purposely tuned to force the models to reason and generalize. It's the whole point of ARC-AGI tests.

Unsurprisingly Gemini 3 pro performs way better on ARC-AGI than 2.5 pro, and unsurprisingly it did much better in pokemon.

The benchmarks, by design, indicate you can mix up the switch puzzle pattern and it will still solve it.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#26
post #4

"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?

If they trained the model to respond to that, then it can respond to that, otherwise it can't necessarily.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#27
post #13
post #8

Earlier quoted context omitted.

There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.

While it is true that model makers are increasingly trying to game benchmarks, it's also true that benchmark-chasing is lowering model quality. GPT 5, 5.1 and 5.2 have been nearly universally panned by almost every class of user, despite being a benchmark monster. In fact, the more OpenAI tries to benchmark-max, the worse their models seem to get.

Hm? 5.1 Thinking is much better than 4o or o3. Just don't use the instant model.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#29
post #4

"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?

If they trained the model to respond to that, then it can respond to that, otherwise it can't necessarily.

I think you got a point here. These companies are injecting a lot of datasets every day into it.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#30

How certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?

You're too kind. Even the CEO of Google retweeted how well Gemini 2.5 did on Pokemon. There is a high chance that now it's explicitly part of the training regime. We kind of need a different kind of game to know how well it generalizes.
Post reply on HN