Live data from Hacker News

Gemini 3

blog.google

601–610 of 1001 posts

Re: Gemini 3

#601

I love it that there's a "Read AI-generated summary" button on their post about their new AI. I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk: > The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you…

Excellent reference Tried to name an AI project at work Electric Monk but too 'controversial'

Had to change to Electric Mentor....

Re: Gemini 3

#602

Earlier quoted context omitted.

What I would do if I was in the position of a large company in this space is to arrange an internal team to create an ARC replica, covering very similar puzzles and use that as part of the training. Ultimately, most benchmarks can be gamed and their real utility is thus short-lived. But I think this is also fair to use any means to beat it.

Is "good at benchmarks instead of real world tasks" really something to optimize for? What does this achieve? Surely people would be initially impressed, try it out, be underwhelmed and then move on. That's not great for Google

If they're memory/reference constrained systems that can't directly "store" every solution, then doing well on benchmarks should result in better real world/reasoning performance, since lack of memorized answer requires understanding.

Like with humans [1], generalized reasoning ability lets you skip the direct storage of that solution, and many many others, completely! You can just synthesize a solution when a problem is presented.

[1] https://www.youtube.com/watch?v=f58kEHx6AQ8

Re: Gemini 3

#603
I don't wan't to shit on the much anticipated G3 model, but I have been using it for a complex single page task and find it underwhelming. Pro 2.5 level, beneath GPT 5.1. Maybe it's launch jitters. It struggles to produce more than 700 lines of code in a single file (aistudio). It struggles to follow instructions. Revisions omit previous gains. I feel cheated! 2.5 Pro has been clearly smarter than everything else for a long time, but now 3 seems not even as good as that, in comparison to the latest releases (5.1 etc). What is going on?

Re: Gemini 3

#604

I love it that there's a "Read AI-generated summary" button on their post about their new AI. I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk: > The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you…

SMBC had a pretty great take on this: https://www.smbc-comics.com/comic/summary

Re: Gemini 3

#605
post #134

Earlier quoted context omitted.

These prediction markets are so ripe for abuse it's unbelievable. People need to realize there are real people on the other side of these bets. Brian Armstong, CEO of Coinbase intentionally altered the outcome of a bet by randomly stating "Bitcoin, Ethereum, blockchain, staking, Web3" at the end of an earnings call. These types of bets shouldn't be allowed.

It’s not really abuse though. These markets aggregate information; when an insider takes one side of a trade, they are selling their information about the true price (probability of the thing happening) to the market (and the price will move accordingly). You’re spot on that people should think of who is on the other side of the trades they’re taking, and be extremely paranoid of being adversely selected. Disallowing…

You don’t get it. Allowing insiders to trade disincentivizes normal people from putting money. Why else is it not allowed in stock market?

Re: Gemini 3

#606

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

Google reports a lower score for Gemini 3 Pro on SWEBench than Claude Sonnet 4.5, which is comparing a top tier model with a smaller one. Very curious to see whether there will be an Opus 4.5 that does even better.

Re: Gemini 3

#607

I am personally impressed by the continued improvement in ARC-AGI-2, where Gemini 3 got 31.1% (vs ChatGPT 5.1's 17.6%). To me this is the kind of problem that does not lend itself well to LLMs - many of the puzzles test the kind of thing that humans intuit because of millions of years of evolution, but these concepts do not necessarily appear in written form (or when they do, it's not clear how they connect to specif…

that looks great, but we all care how it translate to real world problems like programming where it isn't really excelling by 2x.

Re: Gemini 3

#608

Has anyone who is a regular Opus / GPT5-Codex-High / GPT5 Pro user given this model a workout? Each Google release is accompanied by a lot of devrel marketing that sounds impressive but whenever I put the hours into eval myself it comes up lacking. Would love to hear that it replaces another frontier model for someone who is not already bought into the Gemini ecosystem.

I gave it a spin with instructions that worked great with gpt-5-codex (5.1 regressed a lot so I do not even compare to it). Code quality was fine for my very limited tests but I was disappointed with instruction following. I tried few tricks but I wasn't able to convince it to first present plan before starting implementation. I have instructions describing that it should first do exploration (where it tried to disco…

just say "don't code yet" at the end. I never use plan mode because plan mode is just a prompt anyways.

Re: Gemini 3

#609
post #595

Earlier quoted context omitted.

> The ability to add context via a local apps integration into OS level resources is big Good point. I can see why integrated support for local filesystem tools would be useful, even though I prefer manually uploading specific files to avoid polluting the context with irrelevant info. > Its own icon that I can CMD-TAB to is so much nicer Fair enough. I personally prefer Firefox's tab organization to my OS's window or…

> Good point. I can see why integrated support for local filesystem tools would be useful, even though I prefer manually uploading specific files to avoid polluting the context with irrelevant info. Access to OS level resources != context pollution. You still have control, just more direct and less manual. > The ones I've listed are open source and auditable. Yeah I don't plan on spending who knows how much time audi…

> If you are using it because you think it's Open Source I suggest you stop.

I did not know that. Thank you very much for the correction. I guess I have some keys to revoke now.

Re: Gemini 3

#610

Earlier quoted context omitted.

The subtle "wiggle" animation that the second hand makes after moving doesn't fire when it hits 12. Literally unwatchable.

The Swiss and German railway clocks actually work the same way and stop for (half a?) second while the minute handle progresses. https://youtu.be/wejbVtj4YR0

The video shows closer to 2 seconds for it to finally throw itself over in what could only be described as a "Thunk". I figured it would be a little more smooth.
Post reply on HN