Live data from Hacker News

Gemini 3.5 Flash

blog.google

471–480 of 692 posts

Re: Gemini 3.5 Flash

#471

Wow at the price hike. Still I think in the long run the Chinese will win if they're able to produce hardware comparable to Nvidia.

Aren't China also allowed to purchase Nvidia GPUs now too?

Up to the H200 iirc, but they haven't made a purchase yet afaik. The experts in such things believe if they do make a purchase, it will be a token one. Xi is pushing hard for indigenous production, not becoming "hooked" to American Ai chips like some (not so bright people) think we can cause to happen.

Re: Gemini 3.5 Flash

#472
post #139

Earlier quoted context omitted.

If you don't need SOTA or near SOTA there are plenty of dirt cheap models, just look at Gemma 4 31B on Openrouter.

For all of the use cases being hyped you really do, and you actually need something much better than the SOTA models to do what we are being told can be done. The small models are useful for small things like summarizing text or search but not much else.

Yeah a lot of AI hype is look at the amazing new thing our new model can do! Like Google at this event. But when pressed about its pricing reality the answer is “use a worse cheaper model”?? Real convincing argument there

Re: Gemini 3.5 Flash

#473
post #131

The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...

I wonder if they added all these unrequested details as an Easter-egg or something? (Since they must be aware of your test by now).

Re: Gemini 3.5 Flash

#475
post #27

> Create animated SVG of a frog on a boat rowing through jungle river. Single page self contained HTML page with SVG 3.5 Flash: Thinking Medium - 7516 tokens https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0... 3.5 Flash: Thinking High - 7280 tokens https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc165... 3.1 Pro - 28,258 tokens https://gistpreview.github.io/?6bf3da2f80487608b9525bce53018... Though…

It’s shocking how much better 3.1 is than 3.5 flash

The benchmarks used don’t really give a full story

Re: Gemini 3.5 Flash

#476
Taking into account that this is a flash model, it's a strong release. It's very fast and frontier-ish for the price.

Raw intelligence is high for a flash model. But Google's problem has always been productization and tool use, whereas raw intelligence is always competitive. It does not look like they solved that with this release -- in fact, their tool use delta (the improvement in scores when given arbitrary tools and a harness) has actually regressed from some previous models.

Data at https://gertlabs.com/rankings

Re: Gemini 3.5 Flash

#477

For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…

Do you have similar math for the flash-lite variant of the models? I'd be curious. Based on my testing / benchmark i think it's around the 100-120B mark.

With the Pro variant being around 600B - 800B

My testing is comparing it's performance / output to other models in the same size range, so not as scientific as yours.

Re: Gemini 3.5 Flash

#479

Earlier quoted context omitted.

Then why haven't they reported any profits using GAAP (generally accepted accounting principles)? They all use ARR which is easily gamed.

They aren't profitable on a GAAP basis and no one claims this. This obsession over profits is misguided. These are hyper growth companies growing at a scale never seen before. It is both deliberate and uncontroversial to invest in growth rather than slowing down to produce profits.

If my retirement money is going to end up invested in these companies, either directly when they IPO or indirectly through compute providers, then I would like to see some proof that they are capable of producing profits. "Trust me bro" just ain't gonna cut it.

Re: Gemini 3.5 Flash

#480

For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…

Tell me more about what your day looks like. What do you think of the LLMOps books from Abi, in case you have read it ? Any other resources you can recommed?
Post reply on HN