Live data from Hacker News

Gemini 3 Flash: Frontier intelligence built for speed

blog.google

451–460 of 609 posts

Re: Gemini 3 Flash: Frontier intelligence built for speed

#451
post #422

Earlier quoted context omitted.

It's weird they don't document this stuff. Like understanding things like tool call latency and time to first token is extremely important in application development.

Humans often answer with fluff like "That's a good question, thanks for asking that, [fluff, fluff, fluff]" to give themselves more breathing room until the first 'token' of their real answer. I wonder if any LLM are doing stuff like that for latency hiding?

I don't think the models are doing this, time to first token is more of a hardware thing. But people writing agents are definitely doing this, particularly in voice it's worth it to use a smaller local llm to handle the acknowledgment before handing it off.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#452
post #423

Earlier quoted context omitted.

I mean they’re trying to outdo google. So they need to do that.

Until recently, Google was the underdog in the LLM race and OpenAI was the reigning champion. How quickly perceptions shift!

I just want a deepseek moment for an open weights model fast enough to use in my app, I hate paying the big guys.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#453
post #409

Earlier quoted context omitted.

Hardware is a factor here. GPUs are necessarily higher latency than TPUs for equivalent compute on equivalent data. There are lots of other factors here, but latency specifically favours TPUs. The only non-TPU fast models I'm aware of are things running on Cerebras can be much faster because of their CPUs, and Grok has a super fast mode, but they have a cheat code of ignoring guardrails and making up their own world…

> GPUs are necessarily higher latency than TPUs for equivalent compute on equivalent data. Where are you getting that? All the citations I've seen say the opposite, eg: > Inference Workloads: NVIDIA GPUs typically offer lower latency for real-time inference tasks, particularly when leveraging features like NVIDIA's TensorRT for optimized model deployment. TPUs may introduce higher latency in dynamic or low-batch-size…

I'm pretty sure xAI exclusively uses Nvidia H100s for Grok inference but I could be wrong. I agree that I don't see why TPUs would necessarily explain latency.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#454
Only if I could figure out how to use it. I have been using Claude Code and enjoy it. I sometimes also try Codex which is also not bad.

Trying to use Gemini cli is such a pain. I bought GDP Premium and configured GCP, setup environment variables, enabled preview features in cli and did all the dance around it and it won't let me use gemini 3. Why the hell I am even trying so hard?

Re: Gemini 3 Flash: Frontier intelligence built for speed

#455
post #424

Earlier quoted context omitted.

> I don't understand how they expected to sustain the advantage against Google's infinite money machine. I ask this question about Nazi Germany. They adopted the Blitkrieg strategy and expanded unsustainably, but it was only a matter of time until powers with infinite resources (US, USSR) put an end to it.

Huh? How did the USSR have infinite resources? They were barely kept afloat by western allied help (especially at the beginning). Remember also how Tsarist Russia was the first power to collapse and get knocked out of the war in WW1, long before the war was over. They did worse than even the proverbial 'Sick Man of Europe', the Ottoman Empire. Not saying that the Nazi strategy was without flaws, of course. But your s…

they had more soldiers to throw into the meat grinder

Re: Gemini 3 Flash: Frontier intelligence built for speed

#456
post #420

Earlier quoted context omitted.

I didn't see anyone claiming any 'science'? Did I miss something?

I guess there's two things I'm still stuck on: 1. What is the purpose of the benchmark? 2. What is the purpose of publicly discussing a benchmark's results but keeping the methodology secret? To me it's in the same spirit as claiming to have defeated alpha zero but refusing to share the game.

1. The purpose of the benchmark is to choose what models I use for my own system(s). This is extremely common practice in AI - I think every company I've worked with doing LLM work in the last 2 years has done this in some form.

2. I discussed that up-thread, but https://github.com/microsoft/private-benchmarking and https://arxiv.org/abs/2403.00393 discuss some further motivation for this if you are interested.

> To me it's in the same spirit as claiming to have defeated alpha zero but refusing to share the game.

This is an odd way of looking at it. There is no "winning" at benchmarks, it's simply that it is a better and more repeatable evaluation than the old "vibe test" that people did in 2024.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#457
post #173
post #170

Earlier quoted context omitted.

this. I don't know any non-tech people who use anything other than chatgpt. On a similar note, I've wondered why Amazon doesn't make a chatgpt-like app with their latest Alexa+ makeover, seems like a missed opportunity. The Alexa app has a feature to talk to the LLM in chat mode, but the overall app is geared towards managing devices.

Most of Europe if full of Gemini ads, my parents use Gemini because it is free and it popped up in YouTube ad before the video Just go outside the bubble plus take a bit older people

Yeah my parents never really cared enough to explore ChatGPT despite hearing about it 10 times a day in news/media for the last few years. But recently my mom started using Google's AI Search mode after first trying it while doing research for house hunting and my dad uses the Gemini app for occasional questions/identifying parts and stuff (he has always loved Google Lens so those sort of interactive multimedia features are the main pull vs plain text chatbot conversations).

They are both Android/Google Search users so all it really took was "sure I guess I'll try that" in response to a nudge from Google. For me personally I have subscriptions to Claude/ChatGPT/Gemini for coding but use Gemini for 90% of chatbot questions. Eventually I'll cancel some of them but will probably keep Gemini regardless because I like having the extra storage with my Google One plan bundle. Google having a pre-existing platform/ecosystem is a huge advantage imo.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#458
post #407

Earlier quoted context omitted.

I have a bunch of private benchmarks I run against new models I'm evaluating. The reason I don't disclose isn't generally that I think an individual person is going to read my post and update the model to include it. Instead it is because if I write "I ask the question X and expect Y" then that data ends up in the train corpus of new LLMs. However, one set of my benchmarks is a more generalized type of test (think a…

Ok, but then your "post" isn't scientific by definition since it cannot be verified. "Post" is in quotes because I don't know what you're trying to but you're implying some sort of public discourse. For fun: https://chatgpt.com/s/t_694361c12cec819185e9850d0cf0c629

As ChatGPT said to you:

> A secret benchmark is: Useful for internal model selection

That's what I'm doing.

Re: Gemini 3 Flash: Frontier intelligence built for speed

#459
post #326

Earlier quoted context omitted.

> Pretty much every person in the first (and second) world is using AI now This sounds like you live in a huge echo chamber. :-(

Depends what you count as AI (just googling makes you use the LLM summary), but also my mother who is really not tech affine loved what google lense can do, after I showed her. Apart from my very old grandmothers, I don't know anyone not using AI.

I'm sort of old but not a grandmother. Not using AI.
Post reply on HN