Live data from Hacker News

Gemini 3

blog.google

391–400 of 1001 posts

Re: Gemini 3

#391
What I'm getting from this thread is that people have their own private benchmarks. It's almost a cottage industry. Maybe someone should crowd source those benchmarks, keep them completely secret, and create a new public benchmark of people's private AGI tests. All they should release for a given model is the final average score.

Re: Gemini 3

#392

Earlier quoted context omitted.

They've poisoned the internet with their monopoly on advertising, the air pollution of the online world, which is an transgression that far outweighs any good they might have done. Much of the negative social effects of being online come from the need to drive more screen time, more engagement, more clicks, and more ad impressions firehosed into the faces of users for sweet, sweet, advertiser money. When Google final…

They're not a moral entity. corporations aren't people. I think a lot of the harms you mentioned are real, but they're a natural consequence of capitalistic profit chasing. Governments are supposed to regulate monopolies and anti-consumer behavior like that. Instead of regulating surveillance capitalism, governments are using it to bypass laws restricting their power. If I were a google investor, I would absolutely w…

Why should the collective of voters be any more of a moral entity than the collective of people who make up a corporation (which you may include its shareholders in if you want)?

It’s perfectly valid to criticize corporations for their actions, regardless of the regulatory environment.

Re: Gemini 3

#393
Tested it on a bug that Claude and ChatGPT Pro struggled with, it nailed it, but only solved it partially (it was about matching data using a bipartite graph). Another task was optimizing a complex SQL script: the deep-thinking mode provided a genuinely nuanced approach using indexes and rewriting parts of the query. ChatGPT Pro had identified more or less the same issues. For frontend development, I think it’s obvious that it’s more powerful than Claude Code, at least in my tests, the UIs it produces are just better. For backend development, it’s good, but I noticed that in Java specifically, it often outputs code that doesn’t compile on the first try, unlike Claude.

Re: Gemini 3

#394
First impression is I'm having a distinctly harder time getting this to stick to instructions as compared to Gemini 2.5

Re: Gemini 3

#396

Earlier quoted context omitted.

Does anyone trust benchmarks at this point? Genuine question. Isn't the scientific consensus that they are broken and poor evaluation tools?

I make my own automated benchmarks

Is there a tool / website that makes this process easy?

Re: Gemini 3

#397

Tested it on a bug that Claude and ChatGPT Pro struggled with, it nailed it, but only solved it partially (it was about matching data using a bipartite graph). Another task was optimizing a complex SQL script: the deep-thinking mode provided a genuinely nuanced approach using indexes and rewriting parts of the query. ChatGPT Pro had identified more or less the same issues. For frontend development, I think it’s obvio…

> it nailed it, but only solved it partially

Hey either it nailed it or it didn't.

Re: Gemini 3

#398

I think I am in this AI fatigue phase. I am past all hype with models, tools and agents and back to problem and solution approach, sometimes code gen with AI , sometimes think and ask for a piece of code. But not offloading to AI and buying all the bs, waiting it to do magic with my codebase.

I think it's fun to see what is not even considered magic anymore today.

It is. But understandably the people who need to push back on what is still magic may get a bit tired.

Re: Gemini 3

#399

Tested it on a bug that Claude and ChatGPT Pro struggled with, it nailed it, but only solved it partially (it was about matching data using a bipartite graph). Another task was optimizing a complex SQL script: the deep-thinking mode provided a genuinely nuanced approach using indexes and rewriting parts of the query. ChatGPT Pro had identified more or less the same issues. For frontend development, I think it’s obvio…

> it nailed it, but only solved it partially Hey either it nailed it or it didn't.

Probably figured out the exact cause of the bug but not how to solve it

Re: Gemini 3

#400
post #361

Earlier quoted context omitted.

True of almost every new technology.

I hesitate to lump this into the "every new technology" bucket. There are few things that exist today that, similar to what GP said, would have been literal voodoo black magic a few years ago. LLMs are pretty singular in a lot of ways, and you can do powerful things with them that were quite literally impossible a few short years ago. One is free to discount that, but it seems more useful to understand them and their…

More people got more value out of iPhone, including financially.
Post reply on HN