My guess is that they've done a lot of tuning to improve diff based code editing. Gemini 2.5 is fantastic at agentic work, but it still is pretty rough around the edges in terms of generating perfectly matching diffs to edit code. It's probably one of the very few issues with the model. Luckily, aider tracks this. They measure the old gemini 2.5 generating proper diffs 92% of the time. I bet this goes up to ~95-98% h…
Gemini 2.5 Pro Preview
101–110 of 728 posts
Re: Gemini 2.5 Pro Preview
#102Re: Gemini 2.5 Pro Preview
#103Usually don’t believe the benchmarks but first in web dev arena specifically is crazy. That one has been Claude for so long, which tracks in my experience
Re: Gemini 2.5 Pro Preview
#104Is there anything like Claude code for other models such as gemini?
Re: Gemini 2.5 Pro Preview
#105I bet they kept training on coding, made everything worse on the way, and tried to hide it under the rug because of the sunk costs.
Re: Gemini 2.5 Pro Preview
#106I don't know if I'm doing something wrong, but every time I ask gemini 2.5 for code it outputs SO MANY comments. An exaggerated amount of comments. Sections comments, step comments, block comments, inline comments, all the gang.
Re: Gemini 2.5 Pro Preview
#107I don't understand what I'm doing wrong.. it seems like everyone is saying Gemini is better, but I've compared dozens of examples from my work, and Grok has always produced better results.
Re: Gemini 2.5 Pro Preview
#108Why can't they just use version numbers instead of this "new preview" stuff? E.g. call it Gemini Pro 2.5.1.
Re: Gemini 2.5 Pro Preview
#109I'm totally lost again! If I use Gemini on the website (gemini.google.com), am I using 2.5 Pro IO edition, or am I using the old one?
Re: Gemini 2.5 Pro Preview
#110Interestingly, when compering benchmarks of Experimental 03-25 [1] and Experimental 05-06 [2] it seems the new version scores slightly lower in everything except on LiveCodeBench. [1] https://storage.googleapis.com/model-cards/documents/gemini-... [2] https://deepmind.google/technologies/gemini/
This should be the top comment. Cherry-picking is hurting this industry. I bet they kept training on coding tasks, made everything worse on the way, and tried to hide it under the rug because of the sunk costs.