Live data from Hacker News

Gemini 3.5 Flash

blog.google

381–390 of 692 posts

Re: Gemini 3.5 Flash

#381

Earlier quoted context omitted.

latest TPU's appear to reach 800tok/s rather than the advertised 300tok/s.

They demoed today 8i running ate 1300 to 1600ish tokens per second. I imagine that is caused by having a single rack serving the model just for the demo.

There's a limit to how much you can "scale" this process, it's linear, but if we did napkin math based on vllm parallel batched streams only lose around ~50% performance compared to single-stream output so doesn't explain the ridicioulusly fast numbers here.

I wish google just came out and told us how large their flash model is, because if it's as big or smaller than gpt-5.4-nano that's the real headline here.

Re: Gemini 3.5 Flash

#382

Earlier quoted context omitted.

Opus 4.7 is smarter than even Gemini 3.1 Pro on nearly every metric, though. You're comparing apples to oranges. Gemini 3.1 Flash is somewhere in the neighborhood between current Haiku and Sonnet, I think? Still a better value than the Anthropic models, I guess, which are quite pricey. Since Gemini 3.5 Flash is raising the price to $1.50/$9.00, it's priced between Haiku and Sonnet. If it outperforms Sonnet, it remain…

>Opus 4.7 is smarter than even Gemini 3.1 Pro on nearly every metric, Outside of coding, claude models are pretty meh. GPT and Gemini are the workhorses of science/math/finance.

And even on coding, they are mostly good at generating new code.

They sure are not at thorough analysis or debugging, etc.

Re: Gemini 3.5 Flash

#383
post #192

Earlier quoted context omitted.

It is insanely profitable though, if you cut out r&d cost, plus the marketing and loss leaders. Don't let them gaslight you. Even anthropic who does not own any hardware still have a big margin providing claude models.

Everything is insanely profitable if you ignore the costs.

They immediately undercut their argument to the point that I'm not sure if they were being sarcastic.

Re: Gemini 3.5 Flash

#384
They also announced Antigravity CLI, which uses Gemini 3.5 by default. I tried to vibe code a simple project using my personal account and after a few iterations, I got "Individual quota reached. Contact your administrator to enable overages. Resets in [7 days]." Really? 7 days? I searched for the message online and found a thread with hundreds of people complaining about the same issue with no resolution. Classic Google.

Re: Gemini 3.5 Flash

#385
The $1.50/$9.00 pricing is a meaningful shift if you've been running Gemini as the "fast iteration" half of a multi-model coding workflow. I've had Claude Code, Codex, and Gemini CLI running side by side and the working split was "Gemini for quick scaffolding and exploration where the cost of being wrong is low, Sonnet for correctness-critical stuff." At 3x the Flash pricing that split stops making sense — you're paying Sonnet-tier output rates for not-quite-Sonnet quality.

For pure chat that's annoying but tolerable. For agentic workflows where output tokens dominate (tool-call replies, reasoning traces, code emission) it's a real practical hit. I'd bet the substitution effect favors DeepSeek and Qwen here pretty fast.

Re: Gemini 3.5 Flash

#386
post #131

The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...

That pelican looks like it's in Miami for a crypto conference.

They're called ClawCons now

Re: Gemini 3.5 Flash

#387

Earlier quoted context omitted.

According to people who have access to Mythos, it is slightly worse than GPT-5.5-xhigh. At least for security tasks. Hold on, I think this claim needs some hard data. Here you go gentlemen: https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...

That claim keeps contradicted hard by other parties, who say Mythos beats 5.5 resoundingly on both autonomous search and discovery and creation of complex exploit chains. There might be a harness difference, but also, this CTF-type benchmark might not capture the capability difference fully.

[dead]

Re: Gemini 3.5 Flash

#388
post #163
post #27

> Create animated SVG of a frog on a boat rowing through jungle river. Single page self contained HTML page with SVG 3.5 Flash: Thinking Medium - 7516 tokens https://gistpreview.github.io/?5c9858fd2057e678b55d563d9bff0... 3.5 Flash: Thinking High - 7280 tokens https://gistpreview.github.io/?1cab3d70064349d08cf5952cdc165... 3.1 Pro - 28,258 tokens https://gistpreview.github.io/?6bf3da2f80487608b9525bce53018... Though…

Here is GPT 5.5 High thinking; I had to add a second follow up prompt "it's not animated though" as the first one was not animated. https://gistpreview.github.io/?557f979c82701862bc26d24f10399...

Why is it fixated on the front perspective? Interesting choice though, because most humans (and seems like other LLMs too) would pick a side perspective

Re: Gemini 3.5 Flash

#389
post #265

I have google ai pro plan and tried antigravity with 3.5 flash but it used up all my quota in two prompts. If that is not a bug then it is seriously unusable.

Yesterday, or the day before, Google lowered the AI Pro quota from 33x standard usage to 4x. From the talk on the Gemini subreddit it's severely lower than before. I'm likely canceling my AI Pro. The update also broke the app for me. Editing a message crashes the app every time. I'm on a Pixel lol

The crunch is real.

- The model is appox 3.3x cost. - The model is realistically almost 5x cost due to token usage - Google has TPUs to run this on (yet the cost) - Google has a lot more security and backup cash compared to all other AI companies, likely even combined (yet the cost)

We can continue moving the goal posts, but it seems we're at a bit of a wall. Costs are increasing, intelligence is improving, but the cost is rising drastically.

You'd think Google of all companies in the mix would be able to sustain lower costs with how integrated they are with TPU, Deepmind and effectively unlimited budget.

Post reply on HN