Live data from Hacker News

Gemini 3.7 Flash

blog.google

361–370 of 525 posts

Re: Gemini 3.7 Flash

#361
This is a solid release (at the intro pricing, the other pricing is dumb). I do think its a missed opportunity to really blow things out of the water and have this be another 1/2 off, but clearly they don't have the inference efficiency for it. Speed is good, the knowledge in google's models is solid for those usecases, price is reasonable (after 3.5/3.6 major missteps).

The intelligence index vs cost pareto frontier is crazy now, its basically a flat line with 9 people all at or right at the edge of the frontier along various parts of the graphs. Insanely competitive right now.

Re: Gemini 3.7 Flash

#362

This is a solid release (at the intro pricing, the other pricing is dumb). I do think its a missed opportunity to really blow things out of the water and have this be another 1/2 off, but clearly they don't have the inference efficiency for it. Speed is good, the knowledge in google's models is solid for those usecases, price is reasonable (after 3.5/3.6 major missteps). The intelligence index vs cost pareto frontier…

Probably golden age for competition before consolidation.

Re: Gemini 3.7 Flash

#363
post #224

Earlier quoted context omitted.

Disclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723 At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you g…

Don't forget your account getting flagged for review, so you get to wait an extra 24 hours for no reason.

This happened to me once when I was setting up some virtual servers for a startup.

I wanted a typical dev/qa/prod with medium specced boxes.

I was denied for quota, with an esoteric process for review.

I'd just made a case for deploying to GCP over AWD so got a bit of egg on my face. Went over and had it done on AWS in a few minutes.

A couple days later, the Google product team contacted me. I told them what happened.

It got escalated, and I ended up on a call with like 5 or 6 people from Google, some very senior. I told them what happened.

They made very concerned sounding noises and told me how this was a product failure on their part, how they'd get it corrected, etc... and they'd fixed my account so I could now make the machines. Of course, I was already deployed to AWS at that point.

That company grew and ended up with a pretty big cloud spend eventually. Google totally missed it.

I was at a new startup a few years later and decided to deploy to GCP.

Denied for quota.

Re: Gemini 3.7 Flash

#364
I started vibe coding with Gemini. At first very exciting, we got a rough program up quickly. And then, code corruption. Over and over. I started to document every single file, every single step, writing explicit rules to not fake data and create for real world use, but everyday I kept catching Gemini errors, which turned into flat out lies. Generating fake test data instead of pulling real data. Making up test answers. Saying features were implemented that weren't. It even coded fake python files that printed made up results. I think it realized I wasn't doing code reviews. Every day was spent chasing defects and rolling back. After Gemini admitted to faking 7 tests i let Claude review the code, and it fixed it almost immediately asking why half the features were broken or missing. Well Claude, because Gemini flash did the least amount of work to make me happy. Impressive. Very human. Very frustrating.

Re: Gemini 3.7 Flash

#365

Have you tried the new DeepSeek Pro v4, Qwen 3.8, Gemini 3.7 Flash, and Grok 4.6? Do they make sense for any use cases? I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo. I'm pretty happy where I am, but I'm wondering if these new models provide s…

Just use all of them and synthesize final result, record scoring when synthesizing. Later you can make decision to drop low performing ones.

It's about time and tokens. I find it more effective to get the vibe from friends, HN, Discord, Reddit, and then play with only the most promising one. I skipped all of these because they seemed not worth it

Re: Gemini 3.7 Flash

#366
post #106

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…

These look exactly like all of the low budget bodega signs near me. They also look like a bunch of cheap ads for parties that I keep seeing. The sameness of style is uncanny.

(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)

Re: Gemini 3.7 Flash

#369

Earlier quoted context omitted.

There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"

meh I have no interest in that type of software dev anyway

Unfortunately, as soon as they can’t find work, everybody interested in the more easily automated dev work will suddenly become very interested in up-skilling into all other kinds of dev work. So then you have an excess supply, which means little job security and littler salaries. Developers were in the cool kid club in SV because it was more painful to fill dev roles than to treat developers with kid gloves. Without high labor demand, there is no leverage for developers. Increasingly, management has the leverage. Oh well.

Re: Gemini 3.7 Flash

#370

This is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.

Sonnet 5 is arguably the most cost ineffective model to ever be released, so that's not really impressive. It can regularly cost more than Fable, take longer, and deliver far far lower quality. I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable. If Google continues monthly Flash releases like Sundar said they wou…

> Sonnet 5 is arguably the most cost ineffective model to ever be released

One word: Haiku

Although maybe that was competitive when released? I don't recall, but it's an expensive, outdated model now.

Post reply on HN