1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?
Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
21–30 of 97 posts
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#22"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#231.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?
I was unclear if this meant that the API was overloaded or if he was on a subscription plan and had hit his limit for the moment. Although I think that the Gemini plans just use weekly limits, so I guess it must be API.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#24Earlier quoted context omitted.
My issue with this is that the LLM could just be roleplaying that it doesn't know.
To test would just need to edit the rom and switch around the solution. Not sure how complicated that is, likely depends on the rom system.
Basically every benchmark worth it's salt uses bespoke problems purposely tuned to force the models to reason and generalize. It's the whole point of ARC-AGI tests.
Unsurprisingly Gemini 3 pro performs way better on ARC-AGI than 2.5 pro, and unsurprisingly it did much better in pokemon.
The benchmarks, by design, indicate you can mix up the switch puzzle pattern and it will still solve it.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#25Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#26"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#27Earlier quoted context omitted.
There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.
While it is true that model makers are increasingly trying to game benchmarks, it's also true that benchmark-chasing is lowering model quality. GPT 5, 5.1 and 5.2 have been nearly universally panned by almost every class of user, despite being a benchmark monster. In fact, the more OpenAI tries to benchmark-max, the worse their models seem to get.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#281.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#29"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?
If they trained the model to respond to that, then it can respond to that, otherwise it can't necessarily.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#30How certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?