Live data from Hacker News

Gemini 3.7 Flash

blog.google

131–140 of 525 posts

Re: Gemini 3.7 Flash

#131

This is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.

Google is not currently in the lead for maximum model capability, but it is still very competitive (or even best) in the multidimensional capability, cost, and speed frontier.

Re: Gemini 3.7 Flash

#132
It's on Google AI Studio, which I use for free when I'm not on computers I control.

It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.

Re: Gemini 3.7 Flash

#133
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .

It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.

Re: Gemini 3.7 Flash

#134

Gemini Flash is one of the best "good-enough" models. I use this type of model daily, for automation and quick development iteration loops. Unfortunately, it's often not strong enough for heavy refactoring and long running development loops.

'Tis a good workhouse, indeed. I hope they give us a 4.0 Pro that can use Flash subagents soon.

Re: Gemini 3.7 Flash

#135
post #124
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!

Harness could be a part of it, but worth noting both the Opus and Gemini 3.7 flash tests were both ran through opencode.

The grok test was ran through the cursor cli agent however.

Re: Gemini 3.7 Flash

#136
post #25

Earlier quoted context omitted.

> over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. That's why DS4 already had a huge price hike announcement.

The inference providers did not raise the prices no? Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.

[deleted]

Re: Gemini 3.7 Flash

#137
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.

I guess I need to try harder. :)

Re: Gemini 3.7 Flash

#139
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothered building or running. I know a lot of this can be fixed with workflows but it still feels like a failing.

GPT-5.6 or Claude models haven't delivered to me non-running code in ages.

Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.

I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.

As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.

I don't think it's a major selling point when every model can do it well and reasonably fast.

That said, eagerly awaiting "pro" and improvements to antigravity.

Re: Gemini 3.7 Flash

#140
post #57

How does it compare to Opus 5.0 and Fable 5 for coding? E.g. in Cursor or OpenCode?

It is not a competitor to those it competes with Sonnet. Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.

I wasn't aware of this. Seems Google is lagging the big 3 (Anthropic, xAI, OpenAI) when it comes to frontier models for programming and hard problem solving.

I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)

Post reply on HN