Wow at the price hike. Still I think in the long run the Chinese will win if they're able to produce hardware comparable to Nvidia.
Aren't China also allowed to purchase Nvidia GPUs now too?
Gemini 3.5 Flash
471–480 of 692 posts
Re: Gemini 3.5 Flash
#472Earlier quoted context omitted.
If you don't need SOTA or near SOTA there are plenty of dirt cheap models, just look at Gemma 4 31B on Openrouter.
For all of the use cases being hyped you really do, and you actually need something much better than the SOTA models to do what we are being told can be done. The small models are useful for small things like summarizing text or search but not much else.
Re: Gemini 3.5 Flash
#473The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...
Re: Gemini 3.5 Flash
#474Re: Gemini 3.5 Flash
#475> Create animated SVG of a frog on a boat rowing through jungle river. Single page self contained HTML page with SVG 3.5 Flash: Thinking Medium - 7516 tokens https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0... 3.5 Flash: Thinking High - 7280 tokens https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc165... 3.1 Pro - 28,258 tokens https://gistpreview.github.io/?6bf3da2f80487608b9525bce53018... Though…
The benchmarks used don’t really give a full story
Re: Gemini 3.5 Flash
#476Raw intelligence is high for a flash model. But Google's problem has always been productization and tool use, whereas raw intelligence is always competitive. It does not look like they solved that with this release -- in fact, their tool use delta (the improvement in scores when given arbitrary tools and a harness) has actually regressed from some previous models.
Data at https://gertlabs.com/rankings
Re: Gemini 3.5 Flash
#477For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…
With the Pro variant being around 600B - 800B
My testing is comparing it's performance / output to other models in the same size range, so not as scientific as yours.
Re: Gemini 3.5 Flash
#478Re: Gemini 3.5 Flash
#479Earlier quoted context omitted.
Then why haven't they reported any profits using GAAP (generally accepted accounting principles)? They all use ARR which is easily gamed.
They aren't profitable on a GAAP basis and no one claims this. This obsession over profits is misguided. These are hyper growth companies growing at a scale never seen before. It is both deliberate and uncontroversial to invest in growth rather than slowing down to produce profits.
Re: Gemini 3.5 Flash
#480For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…