Live data from Hacker News

Gemini 3 Deep Think

blog.google

281–290 of 722 posts

Re: Gemini 3 Deep Think

#281

Earlier quoted context omitted.

> Could it also be that the models are just a lot better than a year ago? No, the proof is in the pudding. After AI we're having higher prices, higher deficits and lower standard of living. Electricity, computers and everything else costs more. "Doing better" can only be justified by that real benchmark. If Gemini 3 DT was better we would have falling prices of electricity and everything else at least until they get…

You might call me crazy, but at least in 2024, consumers spent ~1% less of their income on expenses than 2019[2], which suggests that 2024 is more affordable than 2019. This is from the BLS consumer survey report released in dec[1] [1] https://www.bls.gov/news.release/cesan.nr0.htm [2] https://www.bls.gov/opub/reports/consumer-expenditures/2019/ Prices are never going back to 2019 numbers though

That's an improper analysis.

First off, it's dollar-averaging every category, so it's not "% of income", which varies based on unit income.

Second, I could commit to spending my entire life with constant spending (optionally inflation adjusted, optionally as a % of income), by adusting quality of goods and service I purchase. So the total spending % is not a measure of affordability.

Re: Gemini 3 Deep Think

#282
post #268

Earlier quoted context omitted.

They are using the current models to help develop even smarter models. Each generation of model can help even more for the next generation. I don’t think it’s hyperbolic to say that we may be only a single digit number of years away from the singularity.

> I don’t think it’s hyperbolic to say that we may be only a single digit number of years away from the singularity. We're back to singularity hype, but let's be real: benchmark gains are meaningless in the real world when the primary focus has shifted to gaming the metrics

Ok, here I am living in the real world finding these models have advanced incredibly over the past year for coding.

Benchmaxxing exists, but that’s not the only data point. It’s pretty clear that models are improving quickly in many domains in real world usage.

Re: Gemini 3 Deep Think

#283

Earlier quoted context omitted.

Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)

Would be cool to have a benchmark with actually unsolved math and science questions, although I suspect models are still quite a long way from that level.

Does folding a protein count? How about increasing performance at Go?

Re: Gemini 3 Deep Think

#284

Earlier quoted context omitted.

Here's a good thread over 1+ month, as each model comes out https://bsky.app/profile/pekka.bsky.social/post/3meokmizvt22... tl;dr - Pekka says Arc-AGI-2 is now toast as a benchmark

If you look at the problem space it is easy to see why it's toast, maybe there's intelligence in there, but hardly general.

> maybe there's intelligence in there, but hardly general.

Of course. Just as our human intelligence isn't general.

Re: Gemini 3 Deep Think

#285
post #74

Not trained for agentic workflows yet unfortunately - this looks like it will be fantastic when they have an agent friendly one. Super exciting.

Its really weird how you all are begging to be replaced by llms, you think if agentic workflows get good enough you're going to keep your job? Or not have your salary reduced by 50%? If Agents get good enough it's not going to build some profitable startup for you (or whatever people think they're doing with the llm slot machines) because that implies that anyone else with access to that agent can just copy you, its…

I agree with you and have similar thoughts (maybe, unfortunately for me). I personally know people who outsource not just their work, but also their life to LLMs, and reading their exciting comments makes me feel a mix of cringe, fomo and dread. But what is the engame for me and you likes, when we finally would be evicted from our own craft? Stash money while we still can, watching 'world crash and burn', and then go and try to ascend in some other, not yet automated craft?

Re: Gemini 3 Deep Think

#286

Earlier quoted context omitted.

You can get more spiky with AIs, whereas with human brain we are more hard wired. So maybe we are forced to be more balanced and general whereas AI don't have to.

I suspect the non-spikey part is the more interesting comparison Why is it so easy for me to open the car door, get in, close the door, buckle up. You can do this in the dark and without looking. There are an infinite number of little things like this you think zero about, take near zero energy, yet which are extremely hard for Ai

You are asking a robotics question, not an AI question. Robotics is more and less than AI. Boston Dynamics robots are getting quite near your benchmark.

Re: Gemini 3 Deep Think

#287
post #159

It's a shame that it's not on OpenRouter. I hate platform lock-in, but the top-tier "deep think" models have been increasingly requiring the use of their own platform.

OpenRouter is pretty great but I think litellm does a very good job and it's not a platform middle man, just a python library. That being said, I have tried it with the deep think models. https://docs.litellm.ai/docs/

Part of OpenRouter's appeal to me is precisely that it is a middle man. I don't want to create accounts on every provider, and juggle all the API keys myself. I suppose this increases my exposure, but I trust all these providers and proxies the same (i.e. not at all), so I'm careful about the data I give them to begin with.

Re: Gemini 3 Deep Think

#288

Earlier quoted context omitted.

That's not a long time in the grand scheme of things.

Speak for yourself. Five years is a long time to wait for my plans of world domination.

This concerns me actually. With enough people (n>=2) wanting to achieve world domination, we have a problem.

Re: Gemini 3 Deep Think

#289
post #53

Arc-AGI-2: 84.6% (vs 68.8% for Opus 4.6) Wow. https://blog.google/innovation-and-ai/models-and-research/ge...

https://arcprize.org/leaderboard $13.62 per task - so we need another 5-10 years for the price to run this to become reasonable? But the real question is if they just fit the model to the benchmark.

A grad student hour is probably more expensive…

Re: Gemini 3 Deep Think

#290

Earlier quoted context omitted.

You are fighting straw men here. Any further discussion would be pointless.

Of course, n-1 wasn't good enough but n+1 will be singularity, just two more weeks my dudes, two more week... rinse and repeat ad infinitum

Like I said, pointless strawmanning.

You’ve once again made up a claim of “two more weeks” to argue against even though it’s not something anybody here has claimed.

If you feel the need to make an argument against claims that exist only in your head, maybe you can also keep the argument only in your head too?

Post reply on HN