Live data from Hacker News

Gemini 3 Deep Think

blog.google

361–370 of 722 posts

Re: Gemini 3 Deep Think

#361

Earlier quoted context omitted.

How can you make sure of that? AFAIK, these SOTA models run exclusively on their developers hardware. So any test, any benchmark, anything you do, does leak per definition. Considering the nature of us humans and the typical prisoners dilemma, I don't see how they wouldn't focus on improving benchmarks even when it gets a bit... shady? I tell this as a person who really enjoys AI by the way.

Because the gains from spending time improving the model overall outweigh the gains from spending time individually training on benchmarks. The pelican benchmark is a good example, because it's been representative of models ability to generate SVGs, not just pelicans on bikes.

> Because the gains from spending time improving the model overall outweigh the gains from spending time individually training on benchmarks.

This may not be the case if you just e.g. roll the benchmarks into the general training data, or make running on the benchmarks just another part of the testing pipeline. I.e. improving the model generally and benchmaxing could very conceivably just both be done at the same time, it needn't be one or the other.

I think the right take away is to ignore the specific percentages reported on these tests (they are almost certainly inflated / biased) and always assume cheating is going on. What matters is that (1) the most serious tests aren't saturated, and (2) scores are improving. I.e. even if there is cheating, we can presume this was always the case, and since models couldn't do as well before even when cheating, these are still real improvements.

And obviously what actually matters is performance on real-world tasks.

Re: Gemini 3 Deep Think

#363

Is it me or is the rate of model release is accelerating to an absurd degree? Today we have Gemini 3 Deep Think and GPT 5.3 Codex Spark. Yesterday we had GLM5 and MiniMax M2.5. Five days before that we had Opus 4.6 and GPT 5.3. Then maybe two weeks I think before that we had Kimi K2.5.

I'm having trouble just keeping track of all these different types of models. Is "Gemini 3 Deep Think" even technically a model? From what I've gathered, it is built on top of Gemini 3 Pro, and appears to be adding specific thinking capabilities, more akin to adding subagents than a truly new foundational model like Opus 4.6. Also, I don't understand the comments about Google being behind in agentic workflows. I know…

The term “model” is one of those super overloaded terms. Depending on the conversation it can mean:

- a product (most accurate here imo)

- a specific set of weights in a neural net

- a general architecture or family of architectures (BERT models)

So while you could argue this is a “model” in the broadest sense of the term, it’s probably more descriptive to call it a product. Similarly we call LLMs “language” models even if they can do a lot more than that, for example draw images.

Re: Gemini 3 Deep Think

#364
post #74

Earlier quoted context omitted.

Its really weird how you all are begging to be replaced by llms, you think if agentic workflows get good enough you're going to keep your job? Or not have your salary reduced by 50%? If Agents get good enough it's not going to build some profitable startup for you (or whatever people think they're doing with the llm slot machines) because that implies that anyone else with access to that agent can just copy you, its…

I’m someone who’d like to deploy a lot more workers than I want to manage. Put another way, I’m on the capital side of the conversation. The good news for labor that has experience and creativity is that it just started costing 1/100,000 what it used to to get on that side of the equation.

If LLMs truly cause widespread replacement of labor, you’re screwed just as much as anyone else. If we hit say 40% unemployment do you think people will care you own your home or not? Do you think people will care you have currency or not? The best case outcome will be universal income and a pseudo utopia where everyone does ok. The “bad” scenario is widespread war.

I am one of the “haves” and am not looking forward to the instability this may bring. Literally no one should.

Re: Gemini 3 Deep Think

#365
post #88

Does anyone actually use Gemini 3 now? I cant stand its sleek salesy way of introduction, and it doesnt hold to instructions hard – makes it unapplicable for MECE breakdowns or for writing.

I use it often. Occasionally for quick questions, but mostly for deep research.

Re: Gemini 3 Deep Think

#366
post #280
post #263

Earlier quoted context omitted.

What is the point of comparing performance of these tools to humans? Machines have been able to accomplish specific tasks better than humans since the industrial revolution. Yet we don't ascribe intelligence to a calculator. None of these benchmarks prove these tools are intelligent, let alone generally intelligent. The hubris and grift are exhausting.

What's the point of denying or downplaying that we are seeing amazing and accelerating advancements in areas that many of us thought were impossible?

It can be reasonable to be skeptical that advances on benchmarks may be only weakly or even negatively correlated with advances on real-world tasks. I.e. a huge jump on benchmarks might not be perceptible to 99% of users doing 99% of tasks, or some users might even note degradation on specific tasks. This is especially the case when there is some reason to believe most benchmarks are being gamed.

Real-world use is what matters, in the end. I'd be surprised if a change this large doesn't translate to something noticeable in general, but the skepticism is not unreasonable here.

Re: Gemini 3 Deep Think

#367
post #294
post #175

Earlier quoted context omitted.

I think it is because of the Chinese new year. The Chinese labs like to publish their models arround the Chinese new year, and the US labs do not want to let a DeepSeek R1 (20 January 2025) impact event happen again, so i guess they publish models that are more capable then what they imagine Chinese labs are yet capable of producing.

[flagged]

Please don't because "Lunar New Year" is ambiguous. Many other Asian cultures also have traditional lunar calendars but a different new years day. It's a bit presumptuous to claim that this is the sole "Lunar New Year" celebration.

https://en.wikipedia.org/wiki/Indian_New_Year%27s_days#Calen...

https://en.wikipedia.org/wiki/Islamic_New_Year

https://en.wikipedia.org/wiki/Nowruz

Re: Gemini 3 Deep Think

#368
post #345

It’s incredible how fast these models are getting better. I thought for sure a wall would be hit, but these numbers smashes previous benchmarks. Anyone have any idea what the big unlock that people are finding now?

Companies are optimizing for all the big benchmarks. This is why there is so little correlation between benchmark performance and real world performance now.

Re: Gemini 3 Deep Think

#370

Is it me or is the rate of model release is accelerating to an absurd degree? Today we have Gemini 3 Deep Think and GPT 5.3 Codex Spark. Yesterday we had GLM5 and MiniMax M2.5. Five days before that we had Opus 4.6 and GPT 5.3. Then maybe two weeks I think before that we had Kimi K2.5.

So, yes, for the past couple weeks it has felt that way to me. But it seems to come in fits and starts. Maybe that will stop being the case, but that's how it's felt to me for awhile.
Post reply on HN