Live data from Hacker News

Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

blog.jcz.dev

81–90 of 97 posts

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#82

So after years of being gleefully told that AI will replace all jobs an omniscient state of the art model, with heavy assistance, takes more than two weeks and thousands of dollars in tokens to do what child me did in a few days? Huh.

“And, because AI never got any better or any cheaper after that point, sussmanbaka’s wry observation remained true in perpetuity, forever.” - History, most likely

We will only be able to see where the economic chips lie once all the current money games shake out. It's all a bit obfuscated at the moment.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#83

Earlier quoted context omitted.

I used to think the same until latest agents started adding perfectly fine features to a large existing react app with just basic input (in English) . Most of the jobs require levels of intelligence below that. It's just a matter of time before agents get to that.

It's about the complexity of the task. Front end apps tend do be much less complex and boilerplate-y than backends, hence AI tends to work better.

Training data is quite readily available as well, and the online education for React is immense in volume. Where enterprise backend software tends to be closed source and unavailable, and there's much less good advice online for how to build with say Java or .NET

That said, I still get surprising results from time to time, it just takes a lot more curation and handholding.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#84
post #79
post #12

1.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?

Who is paying for this? Did the streamer get subsidized by Google? (The stream isn't run by Google themselves, is it?)

If you go to the X page linked on the blog, the page owner mentions a “collaboration” with Google Deepmind on this project. It wouldn’t shock me if this just an elaborate advertisement for Gemini.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#85
post #8

Earlier quoted context omitted.

There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.

If they game the pelican benchmark, it’d be pretty obvious. Just try other random, non-realistic things like “a giraffe walking a tightrope”, “a car sitting at a cafe eating a pizza”, etc. If the results are dramatically different, then they gamed it. If they are similar in quality, then they probably didn’t.

[deleted]

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#87

How certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?

Isn't that the point of a new model anyway?

Yes. Sort of.

Just don’t confuse it with a random benchmark!

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#88

So after years of being gleefully told that AI will replace all jobs an omniscient state of the art model, with heavy assistance, takes more than two weeks and thousands of dollars in tokens to do what child me did in a few days? Huh.

Children are incredibly smart. All of this was fantasy 15 years ago. Comments like yours are amazing to me…

True AI is whatever hasn't been invented yet.

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#89
post #4

"Crucially, it tells the agent not to rely on its internal training data (which might be hallucinated or refer to a different version of the game) but to ground its knowledge in what it observes. " Does this even have any effect?

It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad assumptions (a specific example: the puzzle with Farfetch'd in Ilex Forest is completely different in the DS remake of the game, and models love to hallucinate elements from the remake's puzzle if you don't emphasize the need to distinguish hypothesis from things it actually observes).

Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal

#90

How certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?

There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of Pokemon Red and Blue have gotten.)

That said, this writeup itself will probably be scraped and influence Gemini 4.

Post reply on HN