Live data from Hacker News

Gemini 3

blog.google

361–370 of 1001 posts

Re: Gemini 3

#361

Earlier quoted context omitted.

People would have had a heart attack if they saw this 5 years ago for the first time. Now artificial brains are “meh” :)

True of almost every new technology.

I hesitate to lump this into the "every new technology" bucket. There are few things that exist today that, similar to what GP said, would have been literal voodoo black magic a few years ago. LLMs are pretty singular in a lot of ways, and you can do powerful things with them that were quite literally impossible a few short years ago. One is free to discount that, but it seems more useful to understand them and their strengths, and use them where appropriate.

Even tools like Claude Code have only been fully released for six months, and they've already had a pretty dramatic impact on how many developers work.

Re: Gemini 3

#362
post #237

I have my own private benchmarks for reasoning capabilities on complex problems and i test them against SOTA models regularly (professional cases from law and medicine). Anthropic (Sonnet 4.5 Extended Thinking) and OpenAI (Pro Models) get halfway decent results on many cases while Gemini Pro 2.5 struggled (it was overconfident in its initial assumptions). So i ran these benchmarks against Gemini 3 Pro and i'm not imp…

> It seems very US centric in its thinking

I'm not surprised. I'm French and one thing I've consistently seen with Gemini is that it loves to use Title Case (Everything is Capitalized Except the Prepositions) even in French or other languages where there is no such thing. A 100% american thing getting applied to other languages by the sheer power of statistical correlation (and probably being overtrained on USA-centric data). At the very least it makes it easy to tell when someone is just copypasting LLM output into some other website.

Re: Gemini 3

#363
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

I tried it with gpt-5.1 thinking, and it just searched and found a solution online :p

Is there a solution to this exact problem, or to related notions (renewal equation etc.)? Anyway seems like nothing beats training on test

Re: Gemini 3

#364
I have "unlimited" access to both Gemini 2.5 Pro and Claude 4.5 Sonnet through work.

From my experience, both are capable and can solve nearly all the same complex programming requests, but time and time again Gemini spits out reams and reams of code so over engineered, that totally works, but I would never want to have to interact with.

When looking at the code, you can't tell why it looks "gross", but then you ask Claude to do the same task in the same repo (I use Cline, it's just a dropdown change) and the code also works, but there's a lot less of it and it has a more "elegant" feeling to it.

I know that isn't easy to capture in benchmarks, but I hope Gemini 3.0 has improved in this regard

Re: Gemini 3

#365

Earlier quoted context omitted.

Pffft. OpenAI was conceived to be Open, too.

It’s a common pattern for upstarts to embrace openness as a way to differentiate and gain a foothold then become progressively less open once they get bigger. Android is a great example.

Last I checked, Android is still open source (as AOSP) and people can do whatever-the-f-they-want with the source code. Are we defining open differently?

Re: Gemini 3

#366
Static Pelican is boring. First attempt:

Generate SVG animation of following:

1 - There is High fantasy mage tower with a top window a dome

2 - Green goblin come in front of tower with a torch

3 - Grumpy old mage with beard appear in a tower window in high purple hat

4 - Mage sends fireball that burns goblin and all screen is covered in fire.

Camera view must be from behind of goblin back so we basically look at tower in front of us:

https://codepen.io/Runway/pen/WbwOXRO

Re: Gemini 3

#368

Earlier quoted context omitted.

Gemini app != Google search. You're implying they're lying?

And you're implying they're being 100% truthful? Marketing is always somewhere in the middle

Companies cant get away from egregious marketing. See Apple class action lawsuit for Apple Intelligence.

Re: Gemini 3

#369
The problem with experiencing LLM releases nowadays is that it is no longer trivial to understand the differences in their vast intelligences so it takes awhile to really get a handle on what's even going on.

Re: Gemini 3

#370

Earlier quoted context omitted.

No LLM has ever been as good as people said it was. That doesn't mean this one won't be, but it does make it an unlikely bet based on past trends.

There are 8 Google news articles in the top 15 articles on the HN front page right now. Google being able to skip ahead of every other AI company is wild. They just sat back and watched, then decided it was time to body the competition. The DOJ really should break up Google [1]. They have too many incumbent advantages that were already abuse of monopoly power. [1] https://pluralpolicy.com/find-your-legislator/ - call…

2.5 flash and 2.5 Pro were just sitting back and watching?

The problem with Google is that someone had to show them how to make a product out of the thing, which Open AI did.

Then Anthropic taught them to make a more specific product out of there models

In every aspect, they're just playing catch up, and playing me too.

Models are only part of the solution

Post reply on HN