Anyone using this yet? I’m finding it very bad at instruction following vs 3.1. It calls tools it is told shouldn’t, and it loves calling tools. There’s a pretty strong bias towards its training vs system prompt instructions. Google’s release notes say to reduce unnecessary tool calls by reducing thinking, but that feels like it should be orthogonal to me. It definitely has improved a few logic things, like in data v…
Same. Feels very goal oriented. Requires multiple attempts to deter course and means to achieve it. On tool use. Gave it interactive design assignment on Antigravity 2. Failed miserably until I asked to use playwright for testing. And boy did it go with it. Tested hell out of visuals, nailed the solution. On following instruction. Asked Gemini Flash 3.5 to summarize YouTube video (google io developer keynote), a task…
In my testing, the minimal thinking mode hallucinated 2/3 times, which is pretty scary. The other modes weren’t as bad. I don’t have comprehensive data though.