Live data from Hacker News

Gemini 3.7 Flash

blog.google

261–270 of 525 posts

Re: Gemini 3.7 Flash

#261
post #253

Earlier quoted context omitted.

Does it reset at every turn? From my experience in Codex for example, Luna (Max) fills the 256k token window relatively quick. The only thing lowering the context window again is the compaction.

I mean the thinking does not bloat the context window because it gets dropped at the next request.

It doesn't. It's called preserved reasoning and every recent reasoning model does it

Re: Gemini 3.7 Flash

#262
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one. EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo

clicking the add buttons and scrolling the menu is just much better in Opus 5. It feels like an actual website vs a simulation of one.

Re: Gemini 3.7 Flash

#263
post #41

They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…

flash-lite is more of their luna tier competitor but even still not quite there yet, but gemini's dominance on multimodal and image understanding i think really gets downplayed on this site when most people think the only think you can do with LLMs is write code

Ultimately it would track that in the real world, people will want to point cameras at things and get answers.

I pay for ChatGPT and Gemini, and while Sol is a total beast with anything text, it still poisoned my cucumber bed. Which I will be bitter about for at least a few years while the bed recovers. Gemini (even flash) is exceptionally talented at viewing photos and telling you what to do/what it is (and telling me I just misidentified the problem with my cucumbers and spraying off the "bugs" actually just spread the bacteria everywhere.)

Re: Gemini 3.7 Flash

#264

Model card: https://deepmind.google/models/model-cards/gemini-3-7-flash/ Somewhere in the same neighborhood as GPT 5.6 Tera and Sonnet 5, depending on the bench.

So basically Google is 5 weeks behind with a Fast model that is as good (on the bench they picked it's mostly ahead btw) as the models that the two darlings of HN (OpenAI and Anthropic) released five weeks ago.

And yet the entire thread here is people bitching that it's neither 5.6-sol nor Opus or Fable 5.

BTW why are OpenAI and Anthropic even releasing models like terra/luna and Sonnet?

Why? Just why?

Is there a... market?

For you can't have it both ways: either Sonnet and terra/luna make zero sense for Anthropic and OpenAI or Google is a player.

Re: Gemini 3.7 Flash

#265
i just wish google cloud ux was remotely as good as their models. they made some progress with their studio, but then in a true google fashion, product names keep changing (Gemini, Anti-gravity, Vertex, Google AI,...) as well as confusion and complexity for something as simple as registering agy cli with a Google cloud project.

today i wanted to link agy to a google cloud project, for that i had to enable 5 different APIs in google cloud UI, then create a subscription for Gemini Enterprise (whatever that is), then link it to a project, then assign it to a user. and after all that, agy couldn't find the subscription.

the best part: i couldn't cancel the subscription. so i just paid $35 for one month and left it.

Re: Gemini 3.7 Flash

#266

https://artificialanalysis.ai/models/gemini-3-7-flash The selling point for gemini continues to be speed and particularly end-to-end response time.

You can also customize Gemini Flash. It's a niche thing benefitting few, but you can tune gemini-3.7-flash in Google Vertex (now named "Agent Platform"?)

Re: Gemini 3.7 Flash

#267
post #224

Earlier quoted context omitted.

Can you help me understand how it is hard to get an API key from Google? You just head on over to http://aistudio.google.com/api-keys and create a key... not any different from platform.openai.com? Disclaimer: I work in Google so it might be that this link is not publicly well known

Disclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723 At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you g…

> they're serving companies at Google scale first

I think that's actually a very interesting insight that would be helpful for PMs on GCloud to take note of. As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco, forcing (I assume most) of their users into a arduous process of removing components they don't need.

Google AI Studio is one of Google's solutions to this problem, but in typical Google fashion, it's bolted-on without any clear connection in the ecosystem. If you're also using GCloud, it's hard to remember it's even there.

OpenAI's platform, by contrast, is streamlined, easy to use. With Google, I feel like I need to wade through the documentation first before even using the darn thing.

Re: Gemini 3.7 Flash

#268
post #262

Earlier quoted context omitted.

I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one. EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo

clicking the add buttons and scrolling the menu is just much better in Opus 5. It feels like an actual website vs a simulation of one.

i do feel like opus here is the most natural. theres something off about 3.7 and grok while its an improvement feels flat and not complete

i do wonder why gpt sol was not compared here but honestly it's not really known to be the best at UI

a fable 5 comparison would've been also interesting and likely the best.

Re: Gemini 3.7 Flash

#269
post #140

Earlier quoted context omitted.

It is not a competitor to those it competes with Sonnet. Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.

I wasn't aware of this. Seems Google is lagging the big 3 (Anthropic, xAI, OpenAI) when it comes to frontier models for programming and hard problem solving. I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)

"Lagging" is putting it rather lightly. GDM is no longer a frontier lab.

FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.

Re: Gemini 3.7 Flash

#270

Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens. https://deepswe.datacurve.ai > Starting January 1, 2027,…

I practically switched to doing everything with Luna or DeepSeek V4 flash. I haven't feel the need for the more expensive models.

I'd like to try DS4 if Cursor adds it

I'll use it locally too, but we use Cursor for work

Post reply on HN