Earlier quoted context omitted.
I used to think the same until latest agents started adding perfectly fine features to a large existing react app with just basic input (in English) . Most of the jobs require levels of intelligence below that. It's just a matter of time before agents get to that.
It's about the complexity of the task. Front end apps tend do be much less complex and boilerplate-y than backends, hence AI tends to work better.
Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
51–60 of 97 posts
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#52Earlier quoted context omitted.
It's about the complexity of the task. Front end apps tend do be much less complex and boilerplate-y than backends, hence AI tends to work better.
Isn’t frontend more complex? If my task starts with a Figma UI design, how well does a code agent do at generating working code that looks right, and iterate on it (presuming some browser MCP)? Some automated tests seem enough for an genetic loop on backend.
Haven't tried a Figma design, but i built an internal tool entirely via instructions to agent. The kind of work I could easily quote 3 weeks previously.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#53Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#541.88 billion tokens * $12 / 1M tokens (output) suggests a total cost of $22,560 to solve the game with Gemini 3 Pro?
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#55Earlier quoted context omitted.
There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.
> as they do for popular benchmarks or for penguins riding a bike. Citation?
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#56Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#57So after years of being gleefully told that AI will replace all jobs an omniscient state of the art model, with heavy assistance, takes more than two weeks and thousands of dollars in tokens to do what child me did in a few days? Huh.
- History, most likely
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#58Earlier quoted context omitted.
I used to think the same until latest agents started adding perfectly fine features to a large existing react app with just basic input (in English) . Most of the jobs require levels of intelligence below that. It's just a matter of time before agents get to that.
It's about the complexity of the task. Front end apps tend do be much less complex and boilerplate-y than backends, hence AI tends to work better.
Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#59Re: Gemini 3 Pro vs. 2.5 Pro in Pokemon Crystal
#60How certain can we be that these improvements aren't just a result of Gemini 3 Pro pre-training on endless internet writeups of where 2.5 has struggled (and almost certainly what a human would have done instead)? In other words, how much of this improvement is true generalization vs memorization?