Live data from Hacker News

Gemini 3 Deep Think

blog.google

571–580 of 722 posts

Re: Gemini 3 Deep Think

#571
post #148

Earlier quoted context omitted.

Why would they train on that? Why not just hire someone to make a few examples.

I look forward to them trying. I'll know when the pelican riding a bicycle is good but the ocelot riding a skateboard sucks.

Would it not be better to have 100 such tests "Pelican on bicycle", "Tiger on stilts"..., and generate them all for every new model but only release a new one each time. That way you could show progression across all models, attempts at benchmaxxing would be more obvious.

Given the crazy money and vying for supremacy among AI companies right now it does seem naive to belive that no attempt at better pelicans on bicycles is being made. You can argue "but I will know because of the quality of ocelots on skateboards" but without a back catalog of ocelots on skateboards to publish its one datapoint and leaves the AI companies with too much plausible deniability.

The pelicans-on-bicycles is a bit of fun for you (and us!) but it has become a measure of the quality of models so its serious business for them.

There is an assymetry of incentives and high risk you are being their useful idiot. Sorry to be blunt.

Re: Gemini 3 Deep Think

#573
post #288

Earlier quoted context omitted.

This concerns me actually. With enough people (n>=2) wanting to achieve world domination, we have a problem.

It’s not that I want to achieve world domination (imagine how much work that would be!), it’s just that it’s the inevitable path for AI and I’d rather it be me than then next shmuck with a Claude Max subscription.

Don't build your castle in someone else's kingdom.

Re: Gemini 3 Deep Think

#574
post #535

Too bad we can’t use it. Whenever Google releases something, I can never seem to use it in their coding cli product.

You can but only via Gemini Ultra plan which you can buy or Gemini API with early access.

Re: Gemini 3 Deep Think

#576

When will AI come up with a cure / vaccine for the common cold? and then cancer next?

Race for solving baldness :D

Dutasteride already exists for that, been on it almost 10 years soon and it's great. Although if you are already bald it is kind of moot.

Re: Gemini 3 Deep Think

#578

Earlier quoted context omitted.

Feels like an unforced blunder to make the time window so short after going to so much effort and coming up with something so useful.

5 days for Ai is by no mean short! If it can solve it, it would need perhaps 1-2 hours. If it can not, 5 days continuous running would produce gibberish only. We can safely assume that such private models will run inferences entirely on dedicated hardware, sharing with nobody. So if they could not solve the problems, it's not due to any artificial constraint or lack of resources, far from it. The 5 days window, howev…

5 days is short for memetic propagation on social media to reach everyone who has their own harness and agentic setup that wants to have a go.

Re: Gemini 3 Deep Think

#579

Earlier quoted context omitted.

> does leak per definition. As a measure focused solely on fluid intelligence, learning novel tasks and test-time adaptability, ARC-AGI was specifically designed to be resistant to pre-training - for example, unlike many mathematical and programming test questions, ARC-AGI problems don't have first order patterns which can be learned to solve a different ARC-AGI problem. The ARC non-profit foundation has private vers…

> which is why only "ARC-AGI Certified" results using a secret problem set really matter. The 84.6% is certified and that's a pretty big deal. So, I'd agree if this was on the true fully private set, but Google themselves says they test on only the semi-private: > ARC-AGI-2 results are sourced from the ARC Prize website and are ARC Prize Verified. The set reported is v2, semi-private ( https://storage.googleapis.com/…

Particularly for the large organizations at the frontier, the risk-reward does not seem worth it.

Cheating on the benchmark in such a blatantly intentional way would create a large reputational risk for both the org and the researcher personally.

When you're already at the top, why would you do that just for optimizing one benchmark score?

Re: Gemini 3 Deep Think

#580

Earlier quoted context omitted.

An LLM will claim whatever you tell it to claim. (In fact this Hacker News comment is also conscious.) A dog won’t even claim to be a good boy.

My dog wags his tail hard when I ask "hoosagoodboi?". Pretty definitive I'd say.

[deleted]
Post reply on HN