Live data from Hacker News

Gemini 3

blog.google

811–820 of 1001 posts

Re: Gemini 3

#812

Seems to be the first model that one-shots my secret benchmark about nested SQLite and it did it in 30s,

Out of interest. Does it one shot it every time?

Will try again just tried once in the phone a few hours ago, other models were able to do quite a lot but usually missing some stuff this time it managed nested navigation quite well, lot of stuff missing for sure I just tested the basics with the play button in AI studio

Re: Gemini 3

#813
post #742

Earlier quoted context omitted.

Yeah, it is often pointed out as a brilliance in game analysis if a GM makes a move that an engine says is bad and turns out to be good. However, it only happens in very specific positions.

> Yeah, it is often pointed out as a brilliance in game analysis if a GM makes a move that an engine says is bad and turns out to be good. Do you have any links? I haven't seen any such (forget GM, not even Magnus), barring the opponent making mistakes.

Here’s a chess stackexchange of positions that stump engines

https://chess.stackexchange.com/questions/29716/positions-th...

It basically comes down to “ideas that are rare enough that they were never programmed into a chess engine”.

Blockades or positions where no progress is possible are a common theme. Engines will often keep tree searching where a human sees an obvious repeating pattern.

Here’s also an example where 2 engines are playing, and deep mind finds a move that I think would be obvious to most grandmasters, yet stockfish misses it https://youtu.be/lFXJWPhDsSY?si=zaLQR6sWdEJBMbIO

That being said, I’m not sure that this necessarily correlates with brilliancy. There are a few of these that I would probably get in classical time and I’m not a particularly brilliant player.

Re: Gemini 3

#815

Earlier quoted context omitted.

Out of interest. Does it one shot it every time?

Will try again just tried once in the phone a few hours ago, other models were able to do quite a lot but usually missing some stuff this time it managed nested navigation quite well, lot of stuff missing for sure I just tested the basics with the play button in AI studio

It seems to be that first impression that makes all the difference. Especially with the randomness that comes with llms in general. which maybe explains the 'wow this is so much better' vs the 'this is no better than xxx' commments littered throughout this whole parent post.

Re: Gemini 3

#816
A tad bit better, still has the same issues regarding unpacking and understanding complex prompts. I have a test of mine and now it performs a bit better, but still, it has zero understanding what is happening and for why. Gemini is the best of the best model out there, but with complex problems it just goes down the drain :(.

Re: Gemini 3

#817
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

The problem is these models are optimized to solve the benchmarks, not real world problems.

Re: Gemini 3

#818
post #149

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

Using a single custom benchmark as a metric seems pretty unreliable to me. Even at the risk of teaching future AI the answer to your benchmark, I think you should share it here so we can evaluate it. It's entirely possible you are coming to a wrong conclusion.

No, do not share it. The bigger black hole these models are in, the better.

Re: Gemini 3

#819
post #798

Earlier quoted context omitted.

> They are good at transforming one format to another. They are good at boilerplate. You just described 90% of coding

90% of writing code, sure. But most professionnel programmers write code maybe 20% of the time. A lot of the time is spent clarifying requirements and similar stuff.

The more I hear about other developers' work, the more varied it seems. I've had a few different roles, from one programmer in a huge org to lead programmer in a small team, with a few stints of technical expert in-between. For each the kind of work I do most has varied a lot, but it's never been mostly about "clarifying requirements". As a grunt worker I mostly just wrote and tested code. As a lead I spent most time mentoring, reviewing code, or in meetings. These days I spend most of my time debugging issues and staring at graphics debugger captures.

Re: Gemini 3

#820
Oh that corpulent fella with glasses who talks in the video. Look how good mannered he is, he can't hurt anyone. But Google still takes away all your data and you will be forced out of your job.
Post reply on HN