The intelligence index vs cost pareto frontier is crazy now, its basically a flat line with 9 people all at or right at the edge of the frontier along various parts of the graphs. Insanely competitive right now.
Gemini 3.7 Flash
361–370 of 525 posts
Re: Gemini 3.7 Flash
#362This is a solid release (at the intro pricing, the other pricing is dumb). I do think its a missed opportunity to really blow things out of the water and have this be another 1/2 off, but clearly they don't have the inference efficiency for it. Speed is good, the knowledge in google's models is solid for those usecases, price is reasonable (after 3.5/3.6 major missteps). The intelligence index vs cost pareto frontier…
Re: Gemini 3.7 Flash
#363Earlier quoted context omitted.
Disclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723 At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you g…
Don't forget your account getting flagged for review, so you get to wait an extra 24 hours for no reason.
I wanted a typical dev/qa/prod with medium specced boxes.
I was denied for quota, with an esoteric process for review.
I'd just made a case for deploying to GCP over AWD so got a bit of egg on my face. Went over and had it done on AWS in a few minutes.
A couple days later, the Google product team contacted me. I told them what happened.
It got escalated, and I ended up on a call with like 5 or 6 people from Google, some very senior. I told them what happened.
They made very concerned sounding noises and told me how this was a product failure on their part, how they'd get it corrected, etc... and they'd fixed my account so I could now make the machines. Of course, I was already deployed to AWS at that point.
That company grew and ended up with a pretty big cloud spend eventually. Google totally missed it.
I was at a new startup a few years later and decided to deploy to GCP.
Denied for quota.
Re: Gemini 3.7 Flash
#364Re: Gemini 3.7 Flash
#365Have you tried the new DeepSeek Pro v4, Qwen 3.8, Gemini 3.7 Flash, and Grok 4.6? Do they make sense for any use cases? I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo. I'm pretty happy where I am, but I'm wondering if these new models provide s…
Just use all of them and synthesize final result, record scoring when synthesizing. Later you can make decision to drop low performing ones.
Re: Gemini 3.7 Flash
#366Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…
(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)
Re: Gemini 3.7 Flash
#367Re: Gemini 3.7 Flash
#368I am curious to understand who is this model targeted at?
Re: Gemini 3.7 Flash
#369Earlier quoted context omitted.
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
meh I have no interest in that type of software dev anyway
Re: Gemini 3.7 Flash
#370This is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.
Sonnet 5 is arguably the most cost ineffective model to ever be released, so that's not really impressive. It can regularly cost more than Fable, take longer, and deliver far far lower quality. I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable. If Google continues monthly Flash releases like Sundar said they wou…
One word: Haiku
Although maybe that was competitive when released? I don't recall, but it's an expensive, outdated model now.