Live data from Hacker News

Gemini 3.5 Flash

blog.google

621–630 of 692 posts

Re: Gemini 3.5 Flash

#621
post #583

For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…

If two things hold up - 1) this is actually a 2-300B parameter model and 2) this is actually competitive with frontier OpenAI and Anthropic models (and not just benchmaxing), the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. For comparison, DeepS…

> 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes.

I run 2.54 BPW 397B Qwen 3.5 GGUF on a 128G mac studio at 20 tokens/second generation and 200 tokens/second processing. I'm not suggesting it matches the performance of the full BF16 model, but I did run some benchmarks locally and the results were pretty good:

- MMLU: 87.96%

- GPQA diamond: 86.36%

- IfEval: 91.13%

- GSM8k: 92.57%

So I think we have been at the "frontier capabilities at home" for a few months now.

Re: Gemini 3.5 Flash

#622
post #230

Earlier quoted context omitted.

This is a perfect illustration of something I noticed with llm progress. Ask them to improve an svg like this, and it never fixes the missing crossbar or disconnected limbs, it just adds more stuff. In this example they have obviously improved greatly, and it contains a ridiculous amount of detail, but they still to get the basic shape of the frame wrong. It's weird. And the pattern shows up everywhere, try it with a…

When you say "improve an svg like this", how are you imagining setting that workflow up? Are you just feeding them the SVG to iterate on; or are you giving them access to a browser to look at the rendering of the SVG? I ask because: Insofar as the original pelican test is zero-shot, it effectively serves as a way to test for the presence of a kind of "visual imagination" component within the layers of the model, that…

This is also my gripe with a lot of this stuff, always evaluating models on what they can literally oneshot is completely pointless; it's not how anything works, neither for humans nor for scaffolded AIs. I guess it's neat if you want to argue that a certain level of intelligence can "never be achieved" in a single forward pass, but like, so what. No one cares about that, except people who have already decided to be anti AI.

(not that I am in any sense pro AI, but it's just a weird lack of intellectual rigor)

Re: Gemini 3.5 Flash

#623

Earlier quoted context omitted.

They probably never intended to keep serving cheap models. This is a natural way to introduce the squeeze, now that they have people who built services on their API. It makes a lot of sense to have an abstraction layer where the provider doesn't matter. If you are working in Kotlin, Koog is excellent.

I think the big 3 are cartelizing and starting to ratchet up costs. GPT5.5 is not easily distinguishable from 5.1. I would it be shocked if we hit the ceiling and everyone is quietly positioning for the exit.

I don't understand why everyone thinks there is a ceiling below human-level intelligence, when we have an existence proof that human-level intelligence is possible.

Re: Gemini 3.5 Flash

#624

Earlier quoted context omitted.

If you run out of 50% coupons to your local pizza joint, did they double their prices? Does every company double or triple their prices after Black Friday? There’s a pretty significant difference between saying someone tripled their prices, and a temporary promotion ended. It’s even more so the case if someone is using it as an example for raising prices as a trend. I’m 100% in the camp that prices are going up and q…

> If you run out of 50% coupons to your local pizza joint, did they double their prices? Yes. Did they double their msrp? no. They did double their effective price relative to me which is all that matters unless you're doing economic math or something.

The original comment was used as proof of a trend that vendors are raising prices. Would running out of coupons indicate a trend in rising pizza prices?

Re: Gemini 3.5 Flash

#625
post #609

Earlier quoted context omitted.

> the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. That wouldn't surprise me at all actually, models like Qwen3.6-35B are comparable to frontier level models from a year ago and I wouldn't be surprised if we had self-hostable open weight models matching Opus 4.7 in a year. Assuming that Google has one year of advance against Chinese lab isn't far fetched gi…

I think there was a leap around Opus 4/4.1 that hasn't quite been equalled by self hostable models yet. Perhaps full Kimi K2.6 and Deepseek V4 Pro can achieve Opus 4.1 levels (it's hard to compare anyway, benchmarks are largely a game nowadays), but both of these are also north of 1000B parameters and therefore really impractical to run at home for the foreseeable future. It's not yet obvious to me that you can achie…

> It's not yet obvious to me that you can achieve the breakthrough performance of say Opus 4.1/4.5 in a number of parameters you can swing at home.

People used to believe the same about GPT-4, and I'm not convinced this is going to be different this time.

You do need a very big model if you want something that remembers random trivia about everything, but I'm not convinced this is needed to do meaningful work.

Re: Gemini 3.5 Flash

#626

Earlier quoted context omitted.

If Google is actually getting cheaper inference than everyone else with their TPUs, this smells like trouble to me. Maybe serving LLMs at a profit is proving difficult. Or maybe they think because their benchmarks are good they can ramp up the prices. Seems like they don’t have the market share to justify a move like that yet to me.

This is trouble if you're not Google/OpenAI/Anthropic: they're all shifting towards pricing for the economic value of the knowledge work they're aiding. The economic value increases non-linearly as models get more intelligent: being 10% more capable unlocks way more than 10% in downstream value. That's trouble because the non-linear component means at some point their margins will stop primarily defined by the cost o…

Thank you, this is obviously where we're heading. People who think in terms of "will it ever be profitable to sell tokens" are thinking in the wrong framework entirely. The correct framework is "will it be profitable to sell knowledge work", and the answer will clearly be "yes".

Re: Gemini 3.5 Flash

#627
post #565
post #240

Earlier quoted context omitted.

The cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.

> doesn't mean you get to discount the one business units products to the other That depends, if all developers get used to Claude and Codex it will become harder for Google to attract them in the future. They might lose devs in the long term.

Predatory pricing is a great business strategy and all (particularly when countering the competitors predatory pricing - what could go wrong), but that doesn't mean that the gemini-team should account for it as if they're getting the compute cheaper, it just means that they should run a loss.

Re: Gemini 3.5 Flash

#628

Earlier quoted context omitted.

For all of the use cases being hyped you really do, and you actually need something much better than the SOTA models to do what we are being told can be done. The small models are useful for small things like summarizing text or search but not much else.

Yeah a lot of AI hype is look at the amazing new thing our new model can do! Like Google at this event. But when pressed about its pricing reality the answer is “use a worse cheaper model”?? Real convincing argument there

If don't want to spend 1.5$/9$ for the lastest model then yes use a cheaper model, DeepSeek V4 Flash is 0.11$/0.22$ on OpenRouter and it's more capable than the most expensive model a year ago. Models have never been so cheap given their capabilities unless you want to follow the SOTA (where the hype is).

Re: Gemini 3.5 Flash

#629

Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10

It might be temporary pricing given that 3.5 Flash is actually superior to the existing 3.1 Pro in almost all regards, so they're in a bit of a lurch as 3.1 Pro really doesn't make sense given that 3.5 Pro has been delayed a bit.

I let it loose on a f# codebase that I know was pretty optimized but with a few low hanging fruit changes that would have a big impact.

3.1 Pro did NOT find them. 3.5 flash did. Plus one I hadn't thought of that may or may not work (which it also pointed out).

I'm pretty impressed.

Re: Gemini 3.5 Flash

#630

Earlier quoted context omitted.

Arguably nothing even has to change with training for this to be sustainable. Dario has claimed that Anthropic is profitable on a per training run basis. They aren't profitable because they choose to keep investing in increasingly large training runs.

Cut the crap. The value of the firm's operating assets = EBIT(1-t) - Reinvestment You (Anthropic) want that sky-high valuation? Accept reinvestment is part of the equation. If they decide to stop reinvesting, then they are as good as dead. Moreover, they clearly are not re-investing cash flows from operations. Why do you think they are continually raising money? Lmao.

I'm not sure I understand your argument. If you want an exponentially more expensive training run for each iteration, obviously you need to raise investments even if each training run is profitable. Now I'm not saying that's a good idea, or makes sense, but I am saying that "raising money" doesn't disprove neither that each training run makes money, nor that they're re-investing all that money in the next run.

To give a simple example: if each run simply makes a 10% ROI, but you want to spend 2x as much money on the next run, you still need to raise 90% of the previous run's expenses to have enough capital.

Post reply on HN