Live data from Hacker News

Gemini 3

blog.google

571–580 of 1001 posts

Re: Gemini 3

#572

https://www.youtube.com/watch?v=cUbGVH1r_1U Everyone is talking about the release of Gemini 3. The benchmark scores are incredible. But as we know in the AI world, paper stats don't always translate to production performance on all tasks. We decided to put Gemini 3 through its paces on some standard Vision Language Model (VLM) tasks – specifically simple image detection and processing. The result? It struggled where…

Don't self-promote without disclosure.

Re: Gemini 3

#573
post #538

Earlier quoted context omitted.

Genuinely curious here: why is the desktop app so important? I completely understand the appeal of having local and offline applications, but the ChatGPT desktop app doesn't work without an internet connection anyways. Is it just the convenience? Why is a dedicated desktop app so much better than just opening a browser tab or even using a PWA? Also, have you looked into open-webui or Msty or other provider-agnostic L…

I have a few reasons for the preference: (1) The ability to add context via a local apps integration into OS level resources is big. With Claude, eg, I hit Option-SPC which brings up a prompt bar. From there, taking a screenshot that will get sent my prompt is as simple as dragging a bounding box. This is great. Beyond that, I can add my own MCP connectors and give my desktop app direct access to relevant context in…

> The ability to add context via a local apps integration into OS level resources is big

Good point. I can see why integrated support for local filesystem tools would be useful, even though I prefer manually uploading specific files to avoid polluting the context with irrelevant info.

> Its own icon that I can CMD-TAB to is so much nicer

Fair enough. I personally prefer Firefox's tab organization to my OS's window organization, but I can see how separating the LLM into its own window would be helpful.

> having access to my chats for context has been repeatedly valuable to me.

I didn't at all consider this. Point ceded.

> I haven't looked at provider-agnostic apps and, TBH, would be wary of them.

Interesting. Why? Is it security? The ones I've listed are open source and auditable. I'm confident that they won't steal my API keys. Msty has a lot of advanced functionality that I haven't seen in other interfaces like allowing you to compare responses between different LLMs, export the entire conversation to Markdown, and edit the LLM's response to manage context. It also sidesteps the problem of '[provider] doesn't have a desktop app' because you can use any provider API.

Re: Gemini 3

#574
post #487

Here are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/

Considering how many other "pelican riding a bicycle" comments there are in this thread, it would be surprising if this was not already incorporated in the training data. If not now, soon.

Re: Gemini 3

#575
Trained models should be able to use formal tools (for instance a logical solver, a computer?).

Good. That said, I wonder if those models are still LLMs.

Re: Gemini 3

#576
post #551

Earlier quoted context omitted.

I was interested (and slightly disappointed) to read that the knowledge cutoff for Gemini 3 is the same as for Gemini 2.5: January 2025. I wonder why they didn't train it on more recent data. Is it possible they use the same base pre-trained model and just fine-tuned and RL-ed it better (which, of course, is where all the secret sauce training magic is these days anyhow)? That would be odd, especially for a major ver…

The model card says: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge... > This model is not a modification or a fine-tune of a prior model. I'm curious why they decided not to update the training data cutoff date too.

Maybe that date is a rule of thumb for when AI generated content became so widespread that it is likely to have contaminated future data. Given that people have spoofed authentic Reddit users with Markov chains, it probably doesn’t go back nearly far enough.

Re: Gemini 3

#577
post #47

Earlier quoted context omitted.

Every time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?

This is a list of questions and answers that was created by different people. The questions AND the answers are public. If the LLM manages through reasoning OR memory to repeat back the answer then they win. The scores represent the % of correct answers they recalled.

That is not entirely true. At least some of these tests (like HLE and ARC) take steps to keep the evaluation set private so that LLMs can’t just memorize the answers.

You could question how well this works, but it’s not like the answers are just hanging out on the public internet.

Re: Gemini 3

#578

Earlier quoted context omitted.

Perception seems to be one of the main constraints on LLMs that not much progress has been made on. Perhaps not surprising, given perception is something evolution has worked on since the inception of life itself. Likely much, much more expensive computationally than it receives credit for.

I strongly suspect it's a tokenization problem. Text and symbols fit nicely in tokens, but having something like a single "dog leg" token is a tough problem to solve.

I think in this case, tokenization and percpetion are somewhat analogous. I think it is probably the case our current tokenization schemes are really simplistic compared to what nature is working with. If you allow the analogy.

Re: Gemini 3

#579
I love it that there's a "Read AI-generated summary" button on their post about their new AI.

I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk:

> The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you the bother of washing them yourself; video recorders watched tedious television for you, thus saving you the bother of looking at it yourself. Electric Monks believed things for you, thus saving you what was becoming an increasingly onerous task, that of believing all the things the world expected you to believe.

- from "Dirk Gently's Holistic Detective Agency"

Re: Gemini 3

#580

Earlier quoted context omitted.

I think it's fun to see what is not even considered magic anymore today.

People would have had a heart attack if they saw this 5 years ago for the first time. Now artificial brains are “meh” :)

It is anything but "meh".

It scares the absolute shit out of everyone.

It's clear far beyond our little tech world to everyone this is going to collapse our entire economic system, destroy everyone's livelihoods, and put even more firmly into control the oligarchic assholes already running everything and turning the world to shit.

I see it in news, commentary, day to day conversation. People get it's for real this time and there's a very real chance it ends in something like the Terminator except far worse.

Post reply on HN