Live data from Hacker News

Gemini 3.7 Flash

blog.google

161–170 of 525 posts

Re: Gemini 3.7 Flash

#161
post #135
post #124

Earlier quoted context omitted.

I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!

Harness could be a part of it, but worth noting both the Opus and Gemini 3.7 flash tests were both ran through opencode. The grok test was ran through the cursor cli agent however.

why Grok not through `Grok Build` ?

Re: Gemini 3.7 Flash

#162
post #90

Earlier quoted context omitted.

The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost. And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed: * 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs * ~$12B ARR for Gemini. https://newsletter.semianalysis.…

> The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost. Citation needed. also, why can't a massive company do two things?

With all due respect, did you read my comment beyond the first paragraph? It addresses both points, TPU economics/pivot to sales + internal shortages making it hard to train models, to the extent they can be addressed based on public sources.

There are other factors at play, but they're more recent/second-order.

Re: Gemini 3.7 Flash

#163
post #86
post #41

They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…

Matched roughly with Sol on DeepSwe cost per task. Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.

On DeepSwe it's strictly beaten by Luna on max, cost and result.

Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.

Re: Gemini 3.7 Flash

#164
post #9

The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal. Luna is similar, and also 8x cheaper. Source: artificialanalysis The only benefit I can see is the…

DSV4 Flash is in a tier of its own, until at least the price change arrives.

Re: Gemini 3.7 Flash

#165
post #9

The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal. Luna is similar, and also 8x cheaper. Source: artificialanalysis The only benefit I can see is the…

did you try to ingest 1M documents per hour with any provider except GCP with Flash? None work at scale. Deepseek, Luna, Mistral all fail. 1 in 3 requests is a fail. I stopped trying.

The only thing that works at scale is gemini flash.

Re: Gemini 3.7 Flash

#166
post #41

They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…

Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?

Competitors release multpile models and their curves reflect reasoning effort of each single model.

Gemini doesn't have adjustable reasoning effort (at least on the graph) so each of its curves is just one point.

Re: Gemini 3.7 Flash

#167

Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…

Benchmarks mean very little. The difference between Luna and Sol in the real world is massive.

Re: Gemini 3.7 Flash

#169
post #160
post #133

Earlier quoted context omitted.

Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... . It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really n…

Depends on the definition of friction. If someone is in the Google ecosystem, why would they reach out of it.

I already use GCP and Google for work, and getting an API key was so annoying that even I couldn't be bothered after a while of looking around.

Maybe things there have improved some, but when I was looking it was a huge runaround.

Re: Gemini 3.7 Flash

#170

https://artificialanalysis.ai/models/gemini-3-7-flash The selling point for gemini continues to be speed and particularly end-to-end response time.

It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant. Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...

I use this in a customer facing application and Gemini’s speed makes the experience feel much better.

The application isn’t so complicated that you need opus level reasoning or code writing, we need “good enough” data retrieval and processing with natural language queries and the ability to answer follow up questions.

For that Gemini works well for a decent price.

Post reply on HN