Earlier quoted context omitted.
Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?
For what it's worth the source of the data[0] does have 3.7 flash with all 3 reasoning levels. 3.5/3.6 are in fact just the single points though (high reasoning). The datapoint in the announcement screenshot is either med or high, but they're pretty much exactly the same so can't say for certain. [0]: https://deepswe.datacurve.ai/
Gemini 3.7 Flash
331–340 of 525 posts
Re: Gemini 3.7 Flash
#332Well, I think I'll stick to 3.6 for now. Based in a very scientific sample size of exactly one attempt each: https://imgur.com/a/fDOkBDm
Both got the same prompt and example screenshot. I mean, they both suck but that's normal this early in, but the 3.6 version (first screenshot) actually changes the displayed threads depending on selected categories, and the messages of whatever selected thread, as obviously described in the prompt. You might say it more or less does what it should. Both versions have an ugly flash/jerk in the category pane when selecting/deselecting a category, so that's a wash.
The 3.7 version doesn't work at all, i.e. it always shows all threads, and no messages for any of them. I can post new messages in threads but they don't show up, and it doesn't even increase the message counter for the thread. I guess it's a matter of taste but I don't like the look either, while 3.6 actually is in the spirit of the screenshot I added the prompt, using 11k of CSS versus 3.7's 16k. The code is also less, and the backend split into 3 files (instead of just 2 as 3.7 did it), so assuming it sucks in either case, it'll be easier to read and massage.
edit: geez, 3.6 even properly fades/disables the "new message" button when no thread is selected, 3.7 didn't bother which is smart since everything else is broken anyway. Maybe it's better at really complex things, but for simple things, what I'm experimenting with, I already saw enough.
Re: Gemini 3.7 Flash
#333Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…
Edit: oh wow, diffui looks nice!
Re: Gemini 3.7 Flash
#334The "introductory pricing" for this 3.7 Flash model is really weird. It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago! My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht..…
> got a pretty excellent pelican for the first two This suggests you primarily use Safari. While the bike renders, the pelican doesn’t in Chrome and Firefox. Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.
Re: Gemini 3.7 Flash
#335Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…
I practically switched to doing everything with Luna or DeepSeek V4 flash. I haven't feel the need for the more expensive models.
Re: Gemini 3.7 Flash
#336Google has got the IBM disease. Large lumbering enterprise with massive inertia. Where innovators leave as soon as they get a better offer. None of the authors of the seminal "Attention Is All You Need" paper are still at Google. Fast forward a decade and Google will be reduced to hiring the kind of mediocrities who deign to work at IBM and Accenture.
I don’t think the skill set required to write the Vaswani paper is the same as training and shipping frontier models like Gemini so I am not sure why people keep bringing this up.
Google continuing to ship generations of Gemini using its mediocre teams is an existence proof of that
For genuine advances one needs the Hintons and Vasvanis, not yet another bunch of mediocre Kookaid drinking product managers at Apple and Google
Re: Gemini 3.7 Flash
#337Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…
Re: Gemini 3.7 Flash
#338Has anyone noticed that antigravity has been working really well for the last few weeks. Now with this model it should be working much better. Hope the Google AI Pro Subscription can be used to do some real agentic coding now.
I continually don't understand how nobody points out Flash 3.6 being much faster than any other model, and seems 3.7 is even faster still. That by itself is a major selling point.
Re: Gemini 3.7 Flash
#339Earlier quoted context omitted.
Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.
What would a hedge fund want to do on this exactly? It’s too slow for hft and I’m not sure what they would be doing where ms matter but is not hft.
1. HFT doing ass-simple arbitrage where only latency matters 2. More sophisticated slower trading taking in deeper signals
Those are two points along a continuum. If you are reacting to an earnings announcement by having an LLM read the earnings release and listen to the call, getting the results a few seconds earlier lets you get your trade in a few seconds earlier. Just because "not HFT" doesn't mean "completely latency insensitive".
Re: Gemini 3.7 Flash
#340Earlier quoted context omitted.
GPT-5.6 Luna is an insanely powerful model for its price. It's been great for coding workflows where I guide the LLM's hand step by step. It's also insane to see my weekly limit drop by than 2% after an hour of coding ever since the discount. However, I've noticed 2 drawbacks with Luna. Context rot is much more palpable than Terra and Sol. It tends to get confused and go into rabbit holes when it's context gets fille…
Yes it is cheap, but per task DeepSeek v4 Flash is a bit more expensive and lands between Terra and Gemini 3.6 Flash in quality. Closer to Gemini than Terra...
It's so freaking fast, but you gotta tell Fable to watch Deepseek like a hawk or it'll go off the rails.