Earlier quoted context omitted.
How can system be tuned for a specific model? Model is fungible, often one model strictly greater on both quality and price.
Think of it this way: you are at an enterprise business. You have a workflow implemented a year ago that is working just fine. Swapping out the model for a new one changes behavior in unpredictable ways. Eventually, you'll do it once cost is low enough, but it takes serious labor to validate this, so you'll wait a long enough time for Google to make money.
Gemini 3.7 Flash
81–90 of 525 posts
Re: Gemini 3.7 Flash
#821. How the new model performs against the other top models in the same category.
2. The pricing of the new model against the other top models in the same category.
Re: Gemini 3.7 Flash
#83> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply. this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time. sure we used to cling to gemini models in the past, demanding 2.5 mo…
Re: Gemini 3.7 Flash
#84> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply. this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time. sure we used to cling to gemini models in the past, demanding 2.5 mo…
Maybe the business model is to break even on bleeding edge models while making money on the long tail of usage once systems are tuned for a specific model and running in production.
and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.
Re: Gemini 3.7 Flash
#85The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers. I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal. Luna is similar, and also 8x cheaper. Source: artificialanalysis The only benefit I can see is the…
Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...
If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.
In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.
OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.
Re: Gemini 3.7 Flash
#86They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash. I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost. [edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge... more of a Terra than Luna com…
Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.
Re: Gemini 3.7 Flash
#87https://artificialanalysis.ai/models/gemini-3-7-flash The selling point for gemini continues to be speed and particularly end-to-end response time.
Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster
Re: Gemini 3.7 Flash
#88Earlier quoted context omitted.
Models are not fungible, if you're building certain types of products on them.
this is becoming less true with every generation of model. a decent model with a decent harness will determine when the knowledge base is lacking and attempt to fill the holes; thus the good general models can be very easily brought up to speed on niche domains.
Some model are more aggressive by default, some are more verbose by default. To get the result you want for your specific application, you run experiment with prompts and parameters.
Re: Gemini 3.7 Flash
#89> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply. this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time. sure we used to cling to gemini models in the past, demanding 2.5 mo…
> introductory price They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5' > since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then. This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)
Re: Gemini 3.7 Flash
#90I'm really curious about this: the foundational paper behind today's LLMs came from Google, and some of the world's best scientists were at Google. So why are they falling so far behind in the AI race?
And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed:
* 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs
* ~$12B ARR for Gemini.
https://newsletter.semianalysis.com/p/gemini-is-cooked-but-g...Because they compete for the same scarce resource, the result is a resource crunch for the group that's lost: https://www.latimes.com/business/story/2026-05-18/inside-ai-...