Earlier quoted context omitted.
You dont understand the costs involved to run inference at scale Please go run some numbers.The hardware needed to Run Deepseek v4 flash at 20 tps for a single session is nowhere close to what is required to run it at 50tps for 5,000 concurrent sessions. Imagine what it takes to be profitible when running at 150 tps for 30cents per 1mm. You make less than 1k per month and the hardware required to run that cost 10k a…
Yes it is more efficient in $/tok to run at scale than to run just for yourself. Everyone selling Deepseek V4 inference is selling an undifferentiated good. They have run the numbers on how much it costs and are competing against a dozen other outfits also selling undifferentiated open weights tokens. Whatever the dollar cost they face to rent those GPUs will be what they are able to charge in the competitive market.…
Gemini 3.5 Flash
661–670 of 692 posts
Re: Gemini 3.5 Flash
#662Earlier quoted context omitted.
I think it's unreasonable to expect models generate complex stories in single prompt since they trained to be concise, but I tried. This is prompt on top of story with no control buttons request: Now think, plan how to tell this story in a cartoon, make scene outline and then generate SVG animation story for "Three Little Pigs" in self contained HTML page. Just single animation no control buttons. Full prompt in gist…
This was generated locally with Kimi https://gistpreview.github.io/?d55f07c22d54badc8042a7c8b3785...
Re: Gemini 3.5 Flash
#663Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10
To be fair, Gemini 3.1 flash _lite_ supports structured output (guaranteed json), it’s super fast, runs circles around 2.5 flash and costs $0.25/$1.50. I use it _a lot_ and it’s very capable if you just plan correctly. I actually almost exclusively use 3.1 flash lite and 2.5 flash lite (even cheaper) and we have 99.5% accuracy in what we do. That said, I think we’ll see the lite/flash models and the pro models will d…
Re: Gemini 3.5 Flash
#664Earlier quoted context omitted.
To a certain extent, it feels like a Sonnet 3.7 moment. Slightly overeager - you ask for a button color change, you see layout changes, new package dependencies, and the README rewritten from scratch - and not necessarily correctly. When I ask for a pelican on a bike, I want the Platonic ideal of a pelican on a bike, not a vision of an alternative reality in which pelicans created bikes. Though, thinking about it aga…
What is “Sonnet 3.7 moment”?
Re: Gemini 3.5 Flash
#665Anyone using this yet? I’m finding it very bad at instruction following vs 3.1. It calls tools it is told shouldn’t, and it loves calling tools. There’s a pretty strong bias towards its training vs system prompt instructions. Google’s release notes say to reduce unnecessary tool calls by reducing thinking, but that feels like it should be orthogonal to me. It definitely has improved a few logic things, like in data v…
On tool use. Gave it interactive design assignment on Antigravity 2. Failed miserably until I asked to use playwright for testing. And boy did it go with it. Tested hell out of visuals, nailed the solution.
On following instruction. Asked Gemini Flash 3.5 to summarize YouTube video (google io developer keynote), a task that would previously be trivial (use ot often), but it kept hallucinating points and referencing io dev keynote blog posts from several years ago. Multiple attempts, same result even on repeat requests. Almost insistent on validity of information provided, ignoring questions if it had such capability.
Re: Gemini 3.5 Flash
#666Earlier quoted context omitted.
> If you run out of 50% coupons to your local pizza joint, did they double their prices? Yes. Did they double their msrp? no. They did double their effective price relative to me which is all that matters unless you're doing economic math or something.
The original comment was used as proof of a trend that vendors are raising prices. Would running out of coupons indicate a trend in rising pizza prices?
In this case I'd agree that pricing is effectively raised as 10$ > 10$ - 50%, there's no need to complicate it. However this is not even the right metric for this problem, a better one would be total spent / work produced. If all customers spend more money for the same amount of work (adjusted to progress) then clearly the price is increasing. This would be true in this example as well.
Re: Gemini 3.5 Flash
#667Re: Gemini 3.5 Flash
#668"The reason it is echoing back your messages is because gpt-5.4-nano is a fictional model name!"
"Everything is in perfect order! Let's-Go-ready for the next phase, which will connect this durable infrastructure to the user-facing UI!"
It's like they RLed it on thumbs up and downs on ai overview responses and forgot to make it not be a sycophantic echo chamber machine. And like, the thing it built doesn't work because it's not actually in perfect order, but it doesn't seem to be able to figure out what's wrong because everything is clearly remarkably engineered
Re: Gemini 3.5 Flash
#669Earlier quoted context omitted.
I think the big 3 are cartelizing and starting to ratchet up costs. GPT5.5 is not easily distinguishable from 5.1. I would it be shocked if we hit the ceiling and everyone is quietly positioning for the exit.
I don't understand why everyone thinks there is a ceiling below human-level intelligence, when we have an existence proof that human-level intelligence is possible.
Re: Gemini 3.5 Flash
#670Earlier quoted context omitted.
Elon says Opus is 5T (and I would expect he'd know) > It's not that frontier labs can't create a 5T+ parameter model, but they don't have the data to optimize a model of that size. The have plenty if data. They use very large amounts of verifiable synthetic data in (lots in coding and math) cover the gap. Also the frontier labs are paying people to do tasks, tracking the trajectories and training on that. Most of the…
> Elon says Opus is 5T (and I would expect he'd know) Even if he knew, why would anyone expect Elon not to lie about anything? > The have plenty if data. I don't think data is the problem either, but compute is: if you want to train your 5T params model like modern small models are being trained (with a thousands time more training tokens than params), that's an enormous training run.