Earlier quoted context omitted.
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
FYI, developers are rarely given such a rich UX mock.
Gemini 3.7 Flash
201–210 of 525 posts
Re: Gemini 3.7 Flash
#202Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…
I'm curious how much people are manually curating context these days; I'm increasingly feeling for myself that it being auto-managed inside a front-end like claude code is not ideal, and I'd rather have more control over what exact files and pieces of discovery go into a particular prompt, and the ability to more easily "fork" a session and ask asides or make notes/todos in a way that doesn't disrupt or confuse a mor…
On complete, it moves the user story from Doing to Done. I also have a MEMORY.md file that the agent read and writes in the beginning of a new conversation and at the end of our conversation to update stale information. These files are referred to every time I start a new conversation.
Regarding forking, I know Codex has such button underneath each message that lets you fork the whole conversation. I usually do that when I want to sidetrack and discuss something.
I use no SKILLS or commands like /goal. I’ve come a long way with just prompts and markdown files. Its all a different way of encapsulating instructions anyways.
Re: Gemini 3.7 Flash
#203The "introductory pricing" for this 3.7 Flash model is really weird. It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago! My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht..…
> got a pretty excellent pelican for the first two This suggests you primarily use Safari. While the bike renders, the pelican doesn’t in Chrome and Firefox. Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.
Re: Gemini 3.7 Flash
#204If so, now I understand why they didn't want to release this model
Re: Gemini 3.7 Flash
#205Earlier quoted context omitted.
I've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
I am actively using Gemini flash to "translate" what Opus says into human language. I let opus do the design (with my assistance) and implementation, but then the report that Opus writes gets translated by Gemini so that I don't have to waste time to understand it.
Re: Gemini 3.7 Flash
#206Re: Gemini 3.7 Flash
#207I was going to cancel my gemini membership today ..... still going ahead. In my experience, gemini 3.1 pro, 3.5, 3.6 flash constantly lie too much about completing their tasks whereas sol (even though equally dumb) never claims something has been done when it hasn't been.
1) It makes sense to try 3.7 flash before cancelling.
2) Prompting models to be honest is surprisingly effective in my recent experience. But only if they listen to instructions.
Re: Gemini 3.7 Flash
#208Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…
Benchmarks mean very little. The difference between Luna and Sol in the real world is massive.
Re: Gemini 3.7 Flash
#209Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googling" is for me.
Re: Gemini 3.7 Flash
#210Earlier quoted context omitted.
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... . It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really n…
I like using 3.5-flash-lite for doing cheap PDF and Image data extraction stuff. I don't think there is a better bang / buck model right now (3.1 is cheaper but a lot worse).