Earlier quoted context omitted.
The cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.
> doesn't mean you get to discount the one business units products to the other That depends, if all developers get used to Claude and Codex it will become harder for Google to attract them in the future. They might lose devs in the long term.
Gemini 3.5 Flash
631–640 of 692 posts
Re: Gemini 3.5 Flash
#632For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…
If two things hold up - 1) this is actually a 2-300B parameter model and 2) this is actually competitive with frontier OpenAI and Anthropic models (and not just benchmaxing), the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. For comparison, DeepS…
Re: Gemini 3.5 Flash
#633Earlier quoted context omitted.
I think the big 3 are cartelizing and starting to ratchet up costs. GPT5.5 is not easily distinguishable from 5.1. I would it be shocked if we hit the ceiling and everyone is quietly positioning for the exit.
I don't understand why everyone thinks there is a ceiling below human-level intelligence, when we have an existence proof that human-level intelligence is possible.
Re: Gemini 3.5 Flash
#634Earlier quoted context omitted.
When you say "improve an svg like this", how are you imagining setting that workflow up? Are you just feeding them the SVG to iterate on; or are you giving them access to a browser to look at the rendering of the SVG? I ask because: Insofar as the original pelican test is zero-shot, it effectively serves as a way to test for the presence of a kind of "visual imagination" component within the layers of the model, that…
This is also my gripe with a lot of this stuff, always evaluating models on what they can literally oneshot is completely pointless; it's not how anything works, neither for humans nor for scaffolded AIs. I guess it's neat if you want to argue that a certain level of intelligence can "never be achieved" in a single forward pass, but like, so what. No one cares about that, except people who have already decided to be…
Re: Gemini 3.5 Flash
#635Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.
Re: Gemini 3.5 Flash
#636Earlier quoted context omitted.
This understates the cost increase. 3.5 Flash also uses more tokens. artificialanalysis.ai shows these difference to run the whole eval, which I think is more realistic pricing: Gemini 2.5 flash (27 score): $172 (1.0x) Gemini 2.5 pro (35 score): $649 (3.8x) Gemini 3.0 Flash (46 score): $278 (1.6x) Gemini 3.5 Flash (55 score): $1,552 (9.0x or 2.4x compared to 2.5 pro) This is a massive price increase... 5.6x compared…
Gemini 2.0 Flash: $19
Re: Gemini 3.5 Flash
#637For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…
If two things hold up - 1) this is actually a 2-300B parameter model and 2) this is actually competitive with frontier OpenAI and Anthropic models (and not just benchmaxing), the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. For comparison, DeepS…
Re: Gemini 3.5 Flash
#638Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10
This understates the cost increase. 3.5 Flash also uses more tokens. artificialanalysis.ai shows these difference to run the whole eval, which I think is more realistic pricing: Gemini 2.5 flash (27 score): $172 (1.0x) Gemini 2.5 pro (35 score): $649 (3.8x) Gemini 3.0 Flash (46 score): $278 (1.6x) Gemini 3.5 Flash (55 score): $1,552 (9.0x or 2.4x compared to 2.5 pro) This is a massive price increase... 5.6x compared…
Re: Gemini 3.5 Flash
#639Re: Gemini 3.5 Flash
#640The Flash model costs more than the Frontier models. Didn't see that coming.