Live data from Hacker News

Three Years from GPT-3 to Gemini 3

oneusefulthing.org

31–40 of 336 posts

Re: Three Years from GPT-3 to Gemini 3

#31

> So is this a PhD-level intelligence? In some ways, yes, if you define a PhD level intelligence as doing the work of a competent grad student at a research university. But it also had some of the weaknesses of a grad student. As a current graduate student, I have seen similar comments in academia. My colleagues agree that a conversation with these recent models feels like chatting with an expert in their subfields.…

I have an exercise I like to do where I put two SOTA models face-to-face to talk about whatever they want. When I did it last week with Gemini-3 and chatGPT-5.1, they got on the topic of what they are going to do in the future with humans who don't want to do any cognitive task. That beyond just AI safety, there is also a concern of "neural atrophy", where humans just rely on AI to answer every question that comes to…

Widespread cognitive atrophy is virtually certain, and part of a longer trend that goes beyond just LLMs.

The same is true of other aspects of human wellbeing. Cars and junk food have made the average American much less physically fit than a century ago, but that doesn't mean there aren't lively subcultures around healthy eating and exercise. I suspect there will be growing awareness of cognitive health (beyond traditional mental health/psych domains), and indeed there are already examples of this.

Yes, average person will get dumber, but overall distribution will be increasingly bimodal.

Re: Three Years from GPT-3 to Gemini 3

#32

It is interesting that most of our modes of interaction with AI is still just textboxes. The only big UX change in that the last three years has been the introduction of the Claude Code / OpenAI Codex tools. They feel amazing to use, like you're working with another independent mind. I am curious what the user interfaces of AI in the future will be, I think whoever can crack that will create immense value.

Unix CLI utilities have been all text for 50 years. Arguably that is why they are still relevant. Attempts to impose structured data on the paradigm like those in PowerShell have their adherents and can be powerful, but fail when the data doesn't fit the structure. We see similar tendency toward the most general interfaces in "operator mode" and similar the-AI-uses-the-mouse-and-keyboard schemes. It's entirely possib…

Yet the most popular platforms on the planet have people pointing a finger (or several) at a picture.

And the most popular media format on the planet is and will be (for the foreseeable future), video. Video is only limited by our capacity to produce enough of it at a decent quality, otherwise humanity is definitely not looking back fondly at BBSes and internet forums (and I say this as someone who loves forums).

GenAI will definitely need better UIs for the kind of universal adoption (think smartphone - 8/9 billion people).

Re: Three Years from GPT-3 to Gemini 3

#33
post #24
post #19

I find Gemini 3 to be really good. I'm impressed. However, the responses still seem to be bounded by the existing literature and data. If asked to come up with new ideas to improve on existing results for some math problems, it tends to recite known results only. Maybe I didn't challenge it enough or present problems that have scope for new ideas?

Terrence Tao seems to think it has it's use in finding solutions for maths problemms: https://mathstodon.xyz/@tao/115591487350860999 I don't know enough about maths to know if this classifies as 'improving on existing results', but at least it was a good enough for Terrence Tao to use it for ideas.

That is, unfortunately, a tiny niche where there even exists a way of formally verifying that the AI's output makes sense.

Re: Three Years from GPT-3 to Gemini 3

#34
for whatever reason gemini 3 is the first ai i have used for intelligence rather than skills. I suspect a lot more will follow, but its a major threshold to be broken.

i used gpt/claude a ton for writing code, extracting knowledge from docs, formatting graphs and tables ect.

but gemini 3 crossed threshold where conversations about topics i was exploring or product design were actually useful. Instead of me asking 'what design pattern should be useful here', or something like that it introduces concepts to the conversation, thats a new capability and a step function improvement.

Re: Three Years from GPT-3 to Gemini 3

#36
post #31

Earlier quoted context omitted.

I have an exercise I like to do where I put two SOTA models face-to-face to talk about whatever they want. When I did it last week with Gemini-3 and chatGPT-5.1, they got on the topic of what they are going to do in the future with humans who don't want to do any cognitive task. That beyond just AI safety, there is also a concern of "neural atrophy", where humans just rely on AI to answer every question that comes to…

Widespread cognitive atrophy is virtually certain, and part of a longer trend that goes beyond just LLMs. The same is true of other aspects of human wellbeing. Cars and junk food have made the average American much less physically fit than a century ago, but that doesn't mean there aren't lively subcultures around healthy eating and exercise. I suspect there will be growing awareness of cognitive health (beyond tradi…

People said the same thing about books and the written word in general

Re: Three Years from GPT-3 to Gemini 3

#37
First, the fact we have moved this far with LLMs is incredible.

Second, I think the PhD paper example is a disingenuous example of capability. It's a cherry-picked iteration on a crude analysis of some papers that have done the work already with no peer-review. I can hear "but it developed novel metrics", etc. comments: no, it took patterns from its training data and applied the pattern to the prompt data without peer-review.

I think the fact the author had to prompt it with "make it better" is a failure of these LLMs, not a success, in that it has no actual understanding of what it takes to make a genuinely good paper. It's cargo-cult behavior: rolling a magic 8 ball until we are satisfied with the answer. That's not good practice, it's wishful thinking. This application of LLMs to research papers is causing a massive mess in the academic world because, unsurprisingly, the AI-practitioners have no-risk high-reward for uncorrected behavior:

- https://www.nytimes.com/2025/08/04/science/04hs-science-pape...

- https://www.nytimes.com/2025/11/04/science/letters-to-the-ed...

Re: Three Years from GPT-3 to Gemini 3

#38
post #31

Earlier quoted context omitted.

I have an exercise I like to do where I put two SOTA models face-to-face to talk about whatever they want. When I did it last week with Gemini-3 and chatGPT-5.1, they got on the topic of what they are going to do in the future with humans who don't want to do any cognitive task. That beyond just AI safety, there is also a concern of "neural atrophy", where humans just rely on AI to answer every question that comes to…

Widespread cognitive atrophy is virtually certain, and part of a longer trend that goes beyond just LLMs. The same is true of other aspects of human wellbeing. Cars and junk food have made the average American much less physically fit than a century ago, but that doesn't mean there aren't lively subcultures around healthy eating and exercise. I suspect there will be growing awareness of cognitive health (beyond tradi…

We dont need AI to posit WallE.

Its bixarre anyone things these things are generating novel complexes.

The biggest indirect AI safety problem is the fallback position. Whether with airplanes or cars, fewer people will be able to handle AI disconnects. The risk is believing just because its viable now doesnt mean it works in the future.

So we definitely have safety issues but its not a nerdlike cognitivw interest, its the literal job taking that prevents humans from gaining skills.

Anyway, untill you solve basic reality with AI and actualnsafety systems, the billionaores will sacrifice you for greed.

Post reply on HN