The Gemini Flash models makes perfect sense to me coming from a company like Google. Google AI Mode for search is a product I really find useful. It makes sense that Google focused on smaller, faster yet smart enough models that wouldn't break your bank on inference. It plays well into their product ecosystem. Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googl…
Gemini 3.7 Flash
321–330 of 525 posts
Re: Gemini 3.7 Flash
#322They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…
Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?
Re: Gemini 3.7 Flash
#323They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…
Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?
Re: Gemini 3.7 Flash
#324Earlier quoted context omitted.
Matched roughly with Sol on DeepSwe cost per task. Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.
On DeepSwe it's strictly beaten by Luna on max, cost and result. Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.
Kimi K3 is a beast though, just costly.
Re: Gemini 3.7 Flash
#325Earlier quoted context omitted.
I've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
> I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently. This is what I do too.
Re: Gemini 3.7 Flash
#326Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this. Original images: https://image.non.io/neonRamenDesigns.webp Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7 Opus 5 build for comparison: https://html.non.io/neonRamen Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM pr…
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
Re: Gemini 3.7 Flash
#327I don't get it, Google could heavily subsidy their Gemini models to make it more attractive, but they prefer to not do it. I don't know one soul who is using Gemini models to code. Even OpenAI who doesn't have money or capacity is offering their Luna model at $1.2 per 1M/out.
I am not sure they can afford to subsidize Gemini more than they already do.
Re: Gemini 3.7 Flash
#328Re: Gemini 3.7 Flash
#329Earlier quoted context omitted.
Also best at OpenSCAD, seemingly for the same reason, at least in terms of "iterate on a design, comparing visual output to target".
Are you manually rendering previews of its OpenSCAD output to create images for it to review, or do you have a workflow that automates that?
Re: Gemini 3.7 Flash
#330Earlier quoted context omitted.
gemini flash is probably the best model for visual tasks right now. they also make it really easy to ingest videos
Also best at OpenSCAD, seemingly for the same reason, at least in terms of "iterate on a design, comparing visual output to target".