Live data from Hacker News

Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

blog.jcz.dev

31–40 of 97 posts

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#31
post #4

"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?

It will definitely have some effect. Why won't it? Even adding noise into prompts (like saying you will be rewarded $1000 for each correct answer) has some effect.

Whether the 'effect' something implied by the prompt, or even something we can understand, is a totally different question.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#34
post #32

I like the inclusion of the graph at the end to compare progress. It would be cool to compare this directly to competing models (Claude, GPT, etc).

It would unfortunately also need several runs of each to be reliable. There's nothing in TFA to indicate the results shown aren't to a large degree affected by random chance!

(I do think from personal benchmarks that Gemini 3 is better for the reasons stated by the author, but a single run from each is not strong evidence.)

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#35
post #13
post #8

Earlier quoted context omitted.

There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.

While it is true that model makers are increasingly trying to game benchmarks, it's also true that benchmark-chasing is lowering model quality. GPT 5, 5.1 and 5.2 have been nearly universally panned by almost every class of user, despite being a benchmark monster. In fact, the more OpenAI tries to benchmark-max, the worse their models seem to get.

[deleted]

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#36
post #12

1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?

True though I bet the $200 a month plan could do it, maybe a few extra days of downtime when quota was maxed

For how long would it stay $200 of you can rack up 5 figures if usage..

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#37
post #34
post #32

I like the inclusion of the graph at the end to compare progress. It would be cool to compare this directly to competing models (Claude, GPT, etc).

It would unfortunately also need several runs of each to be reliable. There's nothing in TFA to indicate the results shown aren't to a large degree affected by random chance! (I do think from personal benchmarks that Gemini 3 is better for the reasons stated by the author, but a single run from each is not strong evidence.)

TFA says multiple times that the results are affect by random chance

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#38
post #28
post #12

1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?

:/ Damn. That needs to cost 1000x less before people can try it on their own games.

That's an extrapolation to finish the entire game.

If limit your token count to a fraction of 2 billion tokens, you can try it on your own game, and of course have it complete a shorter fraction of the game.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#39

Earlier quoted context omitted.

True though I bet the $200 a month plan could do it, maybe a few extra days of downtime when quota was maxed

For how long would it stay $200 of you can rack up 5 figures if usage..

That is the reason they severely limited Claude Max subscriptions. Some users racked up 1k+ in API equivalent cost per day.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#40
post #37
post #34

Earlier quoted context omitted.

It would unfortunately also need several runs of each to be reliable. There's nothing in TFA to indicate the results shown aren't to a large degree affected by random chance! (I do think from personal benchmarks that Gemini 3 is better for the reasons stated by the author, but a single run from each is not strong evidence.)

TFA says multiple times that the results are affect by random chance

Yes, but recognising that is only the first step. Quantifying the variance is the next step which I miss in the article.
Post reply on HN