Live data from Hacker News

Gemini 3 Deep Think

blog.google

351–360 of 722 posts

Re: Gemini 3 Deep Think

#351
post #263

Earlier quoted context omitted.

The average ARC AGI 2 score for a single human is around 60%. "100% of tasks have been solved by at least 2 humans (many by more) in under 2 attempts. The average test-taker score was 60%." https://arcprize.org/arc-agi/2/

What is the point of comparing performance of these tools to humans? Machines have been able to accomplish specific tasks better than humans since the industrial revolution. Yet we don't ascribe intelligence to a calculator. None of these benchmarks prove these tools are intelligent, let alone generally intelligent. The hubris and grift are exhausting.

The hubris and grift are exhausting.

And moving the goalposts every few months isn't? What evidence of intelligence would satisfy you?

Personally, my biggest unsatisfied requirement is continual-learning capability, but it's clear we aren't too far from seeing that happen.

Re: Gemini 3 Deep Think

#352
post #281

Earlier quoted context omitted.

That's an improper analysis. First off, it's dollar-averaging every category, so it's not "% of income", which varies based on unit income. Second, I could commit to spending my entire life with constant spending (optionally inflation adjusted, optionally as a % of income), by adusting quality of goods and service I purchase. So the total spending % is not a measure of affordability.

Almost everyone lifestyle ratchets, so the handful that actually downgrade their living rather than increase spending would be tiny. This part of a wider trend too, where economic stats don't align with what people are saying. Which is most likley explained by the economic anomaly of the pandemic skewing peoples perceptions.

We have centuries of historical evidence that people really, really don’t like high inflation, and it takes a while & a lot of turmoil for those shocks to work their way through society.

Re: Gemini 3 Deep Think

#353
post #288

Earlier quoted context omitted.

This concerns me actually. With enough people (n>=2) wanting to achieve world domination, we have a problem.

It’s not that I want to achieve world domination (imagine how much work that would be!), it’s just that it’s the inevitable path for AI and I’d rather it be me than then next shmuck with a Claude Max subscription.

I mean everyone with prompt access to the model says these things, but people like Sam and Elon say these things and mean it.

Re: Gemini 3 Deep Think

#354

Earlier quoted context omitted.

What about Kimi and GLM?

These are well behind the general state of the art (1yr or so), though they're arguably the best openly-available models.

Idk man, GLM 5 in my tests matches opus 4.5 which is what, two months old?

Re: Gemini 3 Deep Think

#355
post #112

Earlier quoted context omitted.

No, not every combination. The question is about the specific combination of a pelican on a bicycle. It might be easy to come up with another test, but we're looking at the results from a particular one here.

More likely you would just train for emitting svg for some description of a scene and create training data from raster images.

None of this works if the testers are collaborating with the trainers. The tests ostensibly need to be arms-length from the training. If the trainers ever start over-fitting to the test, the tester would come up with some new test secretly.

Re: Gemini 3 Deep Think

#356

Earlier quoted context omitted.

François Chollet, creator of ARC-AGI, has consistently said that solving the benchmark does not mean we have AGI. It has always been meant as a stepping stone to encourage progress in the correct direction rather than as an indicator of reaching the destination. That's why he is working on ARC-AGI-3 (to be released in a few weeks) and ARC-AGI-4. His definition of reaching AGI, as I understand it, is when it becomes i…

> His definition of reaching AGI, as I understand it, is when it becomes impossible to construct the next version of ARC-AGI because we can no longer find tasks that are feasible for normal humans but unsolved by AI. That is the best definition I've yet to read. If something claims to be conscious and we can't prove it's not, we have no choice but to believe it. Thats said, I'm reminded of the impossible voting tests…

> If something claims to be conscious and we can't prove it's not, we have no choice but to believe it.

https://x.com/aedison/status/1639233873841201153#m

Re: Gemini 3 Deep Think

#357
post #294
post #175

Earlier quoted context omitted.

I think it is because of the Chinese new year. The Chinese labs like to publish their models arround the Chinese new year, and the US labs do not want to let a DeepSeek R1 (20 January 2025) impact event happen again, so i guess they publish models that are more capable then what they imagine Chinese labs are yet capable of producing.

[flagged]

“Lunar New Year” is vague when referring to the holiday as observed by Chinese labs in China. Chinese people don’t call it Lunar New Year or Chinese New Year anyways. They call it Spring Festival (春节).

As it turns out, people in China don’t name their holidays based off of what the laws of New York or California say.

Re: Gemini 3 Deep Think

#358

Earlier quoted context omitted.

I agree. On top of that, in true Google style, basic things just don't work. Any time I upload an attachment, it just fails with something vague like "couldn't process file". Whether that's a simple .MD or .txt with less than 100 lines or a PDF. I tried making a gem today. It just wouldn't let me save it, with some vague error too. I also tried having it read and write stuff to "my stuff" and Google drive. But it wou…

I don't find that at all. At work, we've no access to the API, so we have to force feed a dozen (or more) documents, code and instruction prompts through the web interface upload interface. The only failures I've ever had in well over 300 sessions were due to connectivity issues, not interface failures. Context window blowouts? All the time, but never document upload failures.

Honestly this is as Google product as you can get. Prizes for some, beatings for others.

Re: Gemini 3 Deep Think

#359

I feel like a luddite: unless I am running small local models, I use gemini-3-flash for almost everything: great for tool use, embedded use in applications, and Python agentic libraries, broad knowledge, good built in web search tool, etc. Oh, and it is fast and cheap. I really only use gemini-3-pro occasionally when researching and trying to better understand something. I guess I am not a good customer for super sca…

[deleted]

Re: Gemini 3 Deep Think

#360
post #294

Earlier quoted context omitted.

[flagged]

"Lunar New Year" is perhaps over-general, since there are non-Asian lunar calendars, such as the Hebrew and Islamic calendars. That said, "Lunar New Year" is probably as good a compromise as any, since we have other names for the Hebrew and Islamic New Years.

This all seems like a plot to get everyone worshipping the Roman goddess Luna.
Post reply on HN