Earlier quoted context omitted.
You dont understand the costs involved to run inference at scale Please go run some numbers.The hardware needed to Run Deepseek v4 flash at 20 tps for a single session is nowhere close to what is required to run it at 50tps for 5,000 concurrent sessions. Imagine what it takes to be profitible when running at 150 tps for 30cents per 1mm. You make less than 1k per month and the hardware required to run that cost 10k a…
Yes it is more efficient in $/tok to run at scale than to run just for yourself. Everyone selling Deepseek V4 inference is selling an undifferentiated good. They have run the numbers on how much it costs and are competing against a dozen other outfits also selling undifferentiated open weights tokens. Whatever the dollar cost they face to rent those GPUs will be what they are able to charge in the competitive market.…
Gemini 3.5 Flash
611–620 of 692 posts
Re: Gemini 3.5 Flash
#612Re: Gemini 3.5 Flash
#613Earlier quoted context omitted.
This understates the cost increase. 3.5 Flash also uses more tokens. artificialanalysis.ai shows these difference to run the whole eval, which I think is more realistic pricing: Gemini 2.5 flash (27 score): $172 (1.0x) Gemini 2.5 pro (35 score): $649 (3.8x) Gemini 3.0 Flash (46 score): $278 (1.6x) Gemini 3.5 Flash (55 score): $1,552 (9.0x or 2.4x compared to 2.5 pro) This is a massive price increase... 5.6x compared…
the era of subsidised ai is ending
Re: Gemini 3.5 Flash
#614For those who would like to know the total and active parameter count of this model: even though Google doesn't disclose the model technicals, we can infer them within relatively tight margins based on what we do know. We know they serve the model on TPU 8i, which we have plenty of hard specs for (so we know the key constraints: total memory and bandwidth and compute flops). We can also set a ceiling on the compute c…
If two things hold up - 1) this is actually a 2-300B parameter model and 2) this is actually competitive with frontier OpenAI and Anthropic models (and not just benchmaxing), the implications are pretty big. It would mean you could run "frontier level" performance in one box at home. 300B models at least fit in a single maxed out Mac Studio or a small stack of DGX Sparks or AMD Strix Halo boxes. For comparison, DeepS…
Re: Gemini 3.5 Flash
#615Re: Gemini 3.5 Flash
#616Earlier quoted context omitted.
People complain about them incessantly, but I can almost never get people to actually post receipts. Every provider allows sharing chats, and anyone can share a prompt that reliably produces hallucinations. More often than not, people are using images in responses that go awry. Which is fair, the models are sold as multi-modal, but image analyses is still at gpt-4.0 text-analyses levels. Also knowledge cutoff issues,…
I see hallucinations ALL the time. It's only obvious when you're prompting about a subject you know well. And when I say all the time, I mean it, and this is for Opus 4.7 Adaptive. I often have to say, please do searches and cite sources, as if it doesn't it will confidently give me wrong or outdated information. If you're often asking questions about a topic that's not in your specialist knowledge you won't notice t…
but for research it makes shit up all the time, I asked GPT5.5 to make me a build for Rogue Trader and not only did it use out of date info, it made up a bunch of skills that were NEVER in the game. I attribute that to there not being enough online information in the wikis or whatever but I wish it would just say "I dont know" instead of hallucinating but I know that's not how the tech works.
Re: Gemini 3.5 Flash
#617Earlier quoted context omitted.
Mate why are you so mad at people upset the price trippeled? It's a fair complaint that people built services using the cheaper ones with the expectation future models would be similarly priced. You can avoid 'offloading thinking' while still building ontop of these models
> It's a fair complaint that people built services using the cheaper ones with the expectation future models would be similarly priced Everyone could see this coming from miles away, everyone warned that this would happen again and again and again, and it always got dismissed.
Re: Gemini 3.5 Flash
#618Earlier quoted context omitted.
thus proving ops point
If you run out of 50% coupons to your local pizza joint, did they double their prices? Does every company double or triple their prices after Black Friday? There’s a pretty significant difference between saying someone tripled their prices, and a temporary promotion ended. It’s even more so the case if someone is using it as an example for raising prices as a trend. I’m 100% in the camp that prices are going up and q…
Yes. Did they double their msrp? no. They did double their effective price relative to me which is all that matters unless you're doing economic math or something.