Live data from Hacker News

Gemini 3.5 Flash

blog.google

661–670 of 692 posts

Re: Gemini 3.5 Flash

#661
post #243

Earlier quoted context omitted.

You dont understand the costs involved to run inference at scale Please go run some numbers.The hardware needed to Run Deepseek v4 flash at 20 tps for a single session is nowhere close to what is required to run it at 50tps for 5,000 concurrent sessions. Imagine what it takes to be profitible when running at 150 tps for 30cents per 1mm. You make less than 1k per month and the hardware required to run that cost 10k a…

Yes it is more efficient in $/tok to run at scale than to run just for yourself. Everyone selling Deepseek V4 inference is selling an undifferentiated good. They have run the numbers on how much it costs and are competing against a dozen other outfits also selling undifferentiated open weights tokens. Whatever the dollar cost they face to rent those GPUs will be what they are able to charge in the competitive market.…

Whoever purchased their RAM last month vs this month has the advantage, I suspect.

Re: Gemini 3.5 Flash

#662
post #207

Earlier quoted context omitted.

I think it's unreasonable to expect models generate complex stories in single prompt since they trained to be concise, but I tried. This is prompt on top of story with no control buttons request: Now think, plan how to tell this story in a cartoon, make scene outline and then generate SVG animation story for "Three Little Pigs" in self contained HTML page. Just single animation no control buttons. Full prompt in gist…

This was generated locally with Kimi https://gistpreview.github.io/?d55f07c22d54badc8042a7c8b3785...

What Kimi exactly? What version and quant?

Re: Gemini 3.5 Flash

#663

Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10

To be fair, Gemini 3.1 flash _lite_ supports structured output (guaranteed json), it’s super fast, runs circles around 2.5 flash and costs $0.25/$1.50. I use it _a lot_ and it’s very capable if you just plan correctly. I actually almost exclusively use 3.1 flash lite and 2.5 flash lite (even cheaper) and we have 99.5% accuracy in what we do. That said, I think we’ll see the lite/flash models and the pro models will d…

I think that’s true on divergence. Basically, the only most is living in the frontier, and even that is only temporary. At some point, the frontier advances such that 99% of tasks can use something short of a frontier model and only a very few tasks actually demand frontier performance.

Re: Gemini 3.5 Flash

#664
post #421
post #375

Earlier quoted context omitted.

To a certain extent, it feels like a Sonnet 3.7 moment. Slightly overeager - you ask for a button color change, you see layout changes, new package dependencies, and the README rewritten from scratch - and not necessarily correctly. When I ask for a pelican on a bike, I want the Platonic ideal of a pelican on a bike, not a vision of an alternative reality in which pelicans created bikes. Though, thinking about it aga…

What is “Sonnet 3.7 moment”?

Sonnet 3.7 tried its damnedest but it was just kinda "off".

Re: Gemini 3.5 Flash

#665

Anyone using this yet? I’m finding it very bad at instruction following vs 3.1. It calls tools it is told shouldn’t, and it loves calling tools. There’s a pretty strong bias towards its training vs system prompt instructions. Google’s release notes say to reduce unnecessary tool calls by reducing thinking, but that feels like it should be orthogonal to me. It definitely has improved a few logic things, like in data v…

Same. Feels very goal oriented. Requires multiple attempts to deter course and means to achieve it.

On tool use. Gave it interactive design assignment on Antigravity 2. Failed miserably until I asked to use playwright for testing. And boy did it go with it. Tested hell out of visuals, nailed the solution.

On following instruction. Asked Gemini Flash 3.5 to summarize YouTube video (google io developer keynote), a task that would previously be trivial (use ot often), but it kept hallucinating points and referencing io dev keynote blog posts from several years ago. Multiple attempts, same result even on repeat requests. Almost insistent on validity of information provided, ignoring questions if it had such capability.

Re: Gemini 3.5 Flash

#666

Earlier quoted context omitted.

> If you run out of 50% coupons to your local pizza joint, did they double their prices? Yes. Did they double their msrp? no. They did double their effective price relative to me which is all that matters unless you're doing economic math or something.

The original comment was used as proof of a trend that vendors are raising prices. Would running out of coupons indicate a trend in rising pizza prices?

It depends whether it's me personally who's running out of coupons or the entire supply of coupons is being reduced. If my ability to get the product for the same price is diminished then the price is being effectively raised.

In this case I'd agree that pricing is effectively raised as 10$ > 10$ - 50%, there's no need to complicate it. However this is not even the right metric for this problem, a better one would be total spent / work produced. If all customers spend more money for the same amount of work (adjusted to progress) then clearly the price is increasing. This would be true in this example as well.

Re: Gemini 3.5 Flash

#668
this model is whack. Exclamation marks everywhere, sycophantic - not producing working code on prompts the other models handle fine.

"The reason it is echoing back your messages is because gpt-5.4-nano is a fictional model name!"

"Everything is in perfect order! Let's-Go-ready for the next phase, which will connect this durable infrastructure to the user-facing UI!"

It's like they RLed it on thumbs up and downs on ai overview responses and forgot to make it not be a sycophantic echo chamber machine. And like, the thing it built doesn't work because it's not actually in perfect order, but it doesn't seem to be able to figure out what's wrong because everything is clearly remarkably engineered

Re: Gemini 3.5 Flash

#669
post #623

Earlier quoted context omitted.

I think the big 3 are cartelizing and starting to ratchet up costs. GPT5.5 is not easily distinguishable from 5.1. I would it be shocked if we hit the ceiling and everyone is quietly positioning for the exit.

I don't understand why everyone thinks there is a ceiling below human-level intelligence, when we have an existence proof that human-level intelligence is possible.

This is very napkin math, but the human brain has about 100 trillion parameters. Even the biggest models today top out at 10 trillion parameters. I think it's reasonable to assume that models need to be at least an order of magnitude bigger to capture the complexity of human intelligence, and probably a lot more.

Re: Gemini 3.5 Flash

#670
post #542

Earlier quoted context omitted.

Elon says Opus is 5T (and I would expect he'd know) > It's not that frontier labs can't create a 5T+ parameter model, but they don't have the data to optimize a model of that size. The have plenty if data. They use very large amounts of verifiable synthetic data in (lots in coding and math) cover the gap. Also the frontier labs are paying people to do tasks, tracking the trajectories and training on that. Most of the…

> Elon says Opus is 5T (and I would expect he'd know) Even if he knew, why would anyone expect Elon not to lie about anything? > The have plenty if data. I don't think data is the problem either, but compute is: if you want to train your 5T params model like modern small models are being trained (with a thousands time more training tokens than params), that's an enormous training run.

I mean in general I'm pretty doubtful about things he says, but in this he was comparing Grok and it sort of makes sense in the context: https://x.com/elonmusk/status/2042123561666855235
Post reply on HN