Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

471–480 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#471
post #219

Earlier quoted context omitted.

There are plenty. But it's not the comparison you want to be making. There is too much variability between the number of tokens used for a single response, especially once reasoning models became a thing. And it gets even worse when you put the models into a variable length output loop. You really need to look at the cost per task. artificialanalysis.ai has a good composite score, measures the cost of running all the…

thanks

For reference the above completely depends on what you're using them for. For many tasks, the number of tokens used is consistent within 10~20%.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#472
post #448

Earlier quoted context omitted.

> I start to believe OAI is very much behind Kara Swisher recently compared OpenAI to Netscape.

Ouch. Maybe we'll get some awesome FOSS tech out of its ashes?

We’ll get a bail-out and then a massive data-centre and energy-production build-out.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#473
post #30

This is awesome. No preview release either, which is great to production. They are pushing the prices higher with each release though: API pricing is up to $0.5/M for input and $3/M for output For comparison: Gemini 3.0 Flash: $0.50/M for input and $3.00/M for output Gemini 2.5 Flash: $0.30/M for input and $2.50/M for output Gemini 2.0 Flash: $0.15/M for input and $0.60/M for output Gemini 1.5 Flash: $0.075/M for inp…

is there a website where i can compare openai, anthropic and gemini models on cost/token ?

https://www.helicone.ai/llm-cost

Tried a lot of them and settled on this one, they update instantly on model release and having all models on one page is the best UX.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#475
post #4

Don’t let the “flash” name fool you, this is an amazing model. I have been playing with it for the past few weeks, it’s genuinely my new favorite; it’s so fast and it has such a vast world knowledge that it’s more performant than Claude Opus 4.5 or GPT 5.2 extra high, for a fraction (basically order of magnitude less!!) of the inference time and price

Oh wow - I recently tried 3 Pro preview and it was too slow for me. After reading your comment I ran my product benchmark against 2.5 flash, 2.5 pro and 3.0 flash. The results are better AND the response times have stayed the same. What an insane gain - especially considering the price compared to 2.5 Pro. I'm about to get much better results for 1/3rd of the price. Not sure what magic Google did here, but would love…

May I ask your internal benchmark ? I'm building a new set of benchmarks and testing suite for agentic workflows using deepwalker [0]. How do you design your benchmark suite ? would be really cool if you can give more details.

[0] https://deepwalker.xyz

Re: Gemini 3 Flash: Frontier intelligence built for speed

#476
post #424

Earlier quoted context omitted.

Huh? How did the USSR have infinite resources? They were barely kept afloat by western allied help (especially at the beginning). Remember also how Tsarist Russia was the first power to collapse and get knocked out of the war in WW1, long before the war was over. They did worse than even the proverbial 'Sick Man of Europe', the Ottoman Empire. Not saying that the Nazi strategy was without flaws, of course. But your s…

they had more soldiers to throw into the meat grinder

They also had more soldiers in WW1.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#477
post #422

Earlier quoted context omitted.

Humans often answer with fluff like "That's a good question, thanks for asking that, [fluff, fluff, fluff]" to give themselves more breathing room until the first 'token' of their real answer. I wonder if any LLM are doing stuff like that for latency hiding?

Do humans really do that often? Coming up with all that fluff would keep my brain busy, meaning there's actually no additional breathing room for thinking about an answer.

People who professionally answer questions do that, yes. Eg politicians or press secretaries for companies, or even just your professor taking questions after a talk.

> Coming up with all that fluff would keep my brain busy, meaning there's actually no additional breathing room for thinking about an answer.

It gets a lot easier with practice: your brain caches a few of the typical fluff routines.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#478
post #308

Earlier quoted context omitted.

OpenAI's doom was written when Altman (and Nadella) got greedy, threw away the nonprofit mission, and caused the exodus of talent and funding that created Anthropic. If they had stayed nonprofit the rest of the industry could have consolidated their efforts against Google's juggernaut. I don't understand how they expected to sustain the advantage against Google's infinite money machine. With Waymo Google showed that…

I think their downfall will be the fact that they don't have a "path to AGI" and have been raising investor money on the promise that they do.

I believethere’s also exponential dislike growing for Altman among most AI users, and that impacts how the brand/company is perceived.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#480
post #456

Earlier quoted context omitted.

1. The purpose of the benchmark is to choose what models I use for my own system(s). This is extremely common practice in AI - I think every company I've worked with doing LLM work in the last 2 years has done this in some form. 2. I discussed that up-thread, but https://github.com/microsoft/private-benchmarking and https://arxiv.org/abs/2403.00393 discuss some further motivation for this if you are interested. > To…

I see the potential value of private evaluations. They aren't scientific but you can certainly beat a "vibe test". I don't understand the value of a public post discussing their results beyond maybe entertainment. We have to trust you implicitly and have no way to validate your claims. > There is no "winning" at benchmarks, it's simply that it is a better and more repeatable evaluation than the old "vibe test" that p…

> I don't understand the value of a public post discussing their results beyond maybe entertainment. We have to trust you implicitly and have no way to validate your claims.

In principle, we have ways: if nl's reports consistently predict how public benchmarks will turn out later, they can build up a reputation. Of course, that requires that we follow nl around for a while.

Post reply on HN