Live data from Hacker News

Gemini 3.7 Flash

blog.google

331–340 of 525 posts

Re: Gemini 3.7 Flash

#331

Earlier quoted context omitted.

Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?

For what it's worth the source of the data[0] does have 3.7 flash with all 3 reasoning levels. 3.5/3.6 are in fact just the single points though (high reasoning). The datapoint in the announcement screenshot is either med or high, but they're pretty much exactly the same so can't say for certain. [0]: https://deepswe.datacurve.ai/

It is extremely odd that this model has that crowbar, spending more for worse results at "high".

Re: Gemini 3.7 Flash

#332
Was excited to try it, since I've been use 3.6 Flash in the last few days to make simply experiments/prototypes. My loop is writing a prompt, maybe adding a screenshot of the closest to what I want I have so far, then based on the result I modify/extend the prompt, maybe use another screenshot.

Well, I think I'll stick to 3.6 for now. Based in a very scientific sample size of exactly one attempt each: https://imgur.com/a/fDOkBDm

Both got the same prompt and example screenshot. I mean, they both suck but that's normal this early in, but the 3.6 version (first screenshot) actually changes the displayed threads depending on selected categories, and the messages of whatever selected thread, as obviously described in the prompt. You might say it more or less does what it should. Both versions have an ugly flash/jerk in the category pane when selecting/deselecting a category, so that's a wash.

The 3.7 version doesn't work at all, i.e. it always shows all threads, and no messages for any of them. I can post new messages in threads but they don't show up, and it doesn't even increase the message counter for the thread. I guess it's a matter of taste but I don't like the look either, while 3.6 actually is in the spirit of the screenshot I added the prompt, using 11k of CSS versus 3.7's 16k. The code is also less, and the backend split into 3 files (instead of just 2 as 3.7 did it), so assuming it sucks in either case, it'll be easier to read and massage.

edit: geez, 3.6 even properly fades/disables the "new message" button when no thread is selected, 3.7 didn't bother which is smart since everything else is broken anyway. Maybe it's better at really complex things, but for simple things, what I'm experimenting with, I already saw enough.

Re: Gemini 3.7 Flash

#333
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

What’s the prompt you used for this?

Edit: oh wow, diffui looks nice!

Re: Gemini 3.7 Flash

#334
post #200
post #141

The "introductory pricing" for this 3.7 Flash model is really weird. It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago! My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht..…

> got a pretty excellent pelican for the first two This suggests you primarily use Safari. While the bike renders, the pelican doesn’t in Chrome and Firefox. Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.

Thank you for pointing this out, I was looking at the first like 'Umm Simon... there isn't even a pelican?"

Re: Gemini 3.7 Flash

#335

Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…

I practically switched to doing everything with Luna or DeepSeek V4 flash. I haven't feel the need for the more expensive models.

Of the two, which do you find better?

Re: Gemini 3.7 Flash

#336

Google has got the IBM disease. Large lumbering enterprise with massive inertia. Where innovators leave as soon as they get a better offer. None of the authors of the seminal "Attention Is All You Need" paper are still at Google. Fast forward a decade and Google will be reduced to hiring the kind of mediocrities who deign to work at IBM and Accenture.

I don’t think the skill set required to write the Vaswani paper is the same as training and shipping frontier models like Gemini so I am not sure why people keep bringing this up.

Training, tweaking and shipping frontier models like Gemini is something that a team of mediocre engineers can do.

Google continuing to ship generations of Gemini using its mediocre teams is an existence proof of that

For genuine advances one needs the Hintons and Vasvanis, not yet another bunch of mediocre Kookaid drinking product managers at Apple and Google

Re: Gemini 3.7 Flash

#337
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

FYI Gemini's version is less broken than Opus' in Safari...

Re: Gemini 3.7 Flash

#338

Has anyone noticed that antigravity has been working really well for the last few weeks. Now with this model it should be working much better. Hope the Google AI Pro Subscription can be used to do some real agentic coding now.

I've been really satisfied with it since 3.6, it's been "good enough" for the tasks I'm using and has fast response and very high limits, much higher than Claude Code.

I continually don't understand how nobody points out Flash 3.6 being much faster than any other model, and seems 3.7 is even faster still. That by itself is a major selling point.

Re: Gemini 3.7 Flash

#339

Earlier quoted context omitted.

Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.

What would a hedge fund want to do on this exactly? It’s too slow for hft and I’m not sure what they would be doing where ms matter but is not hft.

It's not like there are only two buckets:

1. HFT doing ass-simple arbitrage where only latency matters 2. More sophisticated slower trading taking in deeper signals

Those are two points along a continuum. If you are reacting to an earnings announcement by having an LLM read the earnings release and listen to the call, getting the results a few seconds earlier lets you get your trade in a few seconds earlier. Just because "not HFT" doesn't mean "completely latency insensitive".

Re: Gemini 3.7 Flash

#340
post #259

Earlier quoted context omitted.

GPT-5.6 Luna is an insanely powerful model for its price. It's been great for coding workflows where I guide the LLM's hand step by step. It's also insane to see my weekly limit drop by than 2% after an hour of coding ever since the discount. However, I've noticed 2 drawbacks with Luna. Context rot is much more palpable than Terra and Sol. It tends to get confused and go into rabbit holes when it's context gets fille…

Yes it is cheap, but per task DeepSeek v4 Flash is a bit more expensive and lands between Terra and Gemini 3.6 Flash in quality. Closer to Gemini than Terra...

Fable orchestrating DeepSeek v4 Flash to implement a plan is my new favorite thing.

It's so freaking fast, but you gotta tell Fable to watch Deepseek like a hawk or it'll go off the rails.

Post reply on HN