Live data from Hacker News

Gemini 3.5 Flash

blog.google

561–570 of 692 posts

Re: Gemini 3.5 Flash

#561

For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…

I like your chain of thought there !

Re: Gemini 3.5 Flash

#563
post #83

The price is crazy. And I guess Gemini 3.5 pro will have the pricing increment, too. 12 x 5 = 60? It seems like google does want us to use Chinese models.

What exactly are you doing with this that you can’t generate $1.50 of value per million tokens?

I sell service. Imagine my users have to pay 4x more for marginal increment just 'cause.

They are more willing to wait though, so Chinese models are pretty attractive right now.

Re: Gemini 3.5 Flash

#564

Earlier quoted context omitted.

Deepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.

Mate why are you so mad at people upset the price trippeled? It's a fair complaint that people built services using the cheaper ones with the expectation future models would be similarly priced. You can avoid 'offloading thinking' while still building ontop of these models

> It's a fair complaint that people built services using the cheaper ones with the expectation future models would be similarly priced

Everyone could see this coming from miles away, everyone warned that this would happen again and again and again, and it always got dismissed.

Re: Gemini 3.5 Flash

#565
post #240

Earlier quoted context omitted.

This is not priced at inference cost. My guess: it's the price at which they make more money than if they rent the TPUs to other companies. The Gemini team has had trouble securing enough TPUs for their user's needs. They struggle with load and their rate limits are really bad. Maybe at a higher price, they have a better chance at getting more TPUs assigned?

The cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.

> doesn't mean you get to discount the one business units products to the other

That depends, if all developers get used to Claude and Codex it will become harder for Google to attract them in the future.

They might lose devs in the long term.

Re: Gemini 3.5 Flash

#566

For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…

We've been really impressed with the performance of ~30B parameter class models and how close they are to the frontier from ~6-12 months ago, which begs the question, are the frontier labs really serving 10T parameter models? Seems unlikely. If these Gemini 3.5 numbers are accurate, then I'd wager GPT 5.5 and Opus 4.7 are a lot smaller than people have speculated, too. It's not that frontier labs can't create a 5T+ p…

It was estimated that Mythos is 10T.

And serving is not training. For distilling you need to train the big models to have something to be distilled.

Re: Gemini 3.5 Flash

#567
Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.

Re: Gemini 3.5 Flash

#568
post #426

Earlier quoted context omitted.

I agree with parent. I'm not sure where your stance is coming from. From what I hear, most enterprise AI deployments are seat-based subscriptions with annual commitments.

Yes, I work at a 50 person startup and even here switching from CC to codex or cursor would be non-trivial for multiple reasons - not just the annual commitment.

I don't doubt you but it's amazing how much easier things get when there's another option at 20% of the price, and that's what's going to happen here if these American companies keep trying to squeeze the prices up.

Re: Gemini 3.5 Flash

#569

Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.

Yeap it russian, but the whole russian sentence doesn't make any sense, just messed words with no meaning at all :)

Re: Gemini 3.5 Flash

#570

Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.

Yeap it russian, but the whole russian sentence doesn't make any sense, just messed words with no meaning at all :)

But the voice and pauses sounds so much real, it's hard to say "it was ai", sounds like a real human
Post reply on HN