Gemini 3
571–580 of 1001 posts
Re: Gemini 3
#572https://www.youtube.com/watch?v=cUbGVH1r_1U Everyone is talking about the release of Gemini 3. The benchmark scores are incredible. But as we know in the AI world, paper stats don't always translate to production performance on all tasks. We decided to put Gemini 3 through its paces on some standard Vision Language Model (VLM) tasks – specifically simple image detection and processing. The result? It struggled where…
Re: Gemini 3
#573Earlier quoted context omitted.
Genuinely curious here: why is the desktop app so important? I completely understand the appeal of having local and offline applications, but the ChatGPT desktop app doesn't work without an internet connection anyways. Is it just the convenience? Why is a dedicated desktop app so much better than just opening a browser tab or even using a PWA? Also, have you looked into open-webui or Msty or other provider-agnostic L…
I have a few reasons for the preference: (1) The ability to add context via a local apps integration into OS level resources is big. With Claude, eg, I hit Option-SPC which brings up a prompt bar. From there, taking a screenshot that will get sent my prompt is as simple as dragging a bounding box. This is great. Beyond that, I can add my own MCP connectors and give my desktop app direct access to relevant context in…
Good point. I can see why integrated support for local filesystem tools would be useful, even though I prefer manually uploading specific files to avoid polluting the context with irrelevant info.
> Its own icon that I can CMD-TAB to is so much nicer
Fair enough. I personally prefer Firefox's tab organization to my OS's window organization, but I can see how separating the LLM into its own window would be helpful.
> having access to my chats for context has been repeatedly valuable to me.
I didn't at all consider this. Point ceded.
> I haven't looked at provider-agnostic apps and, TBH, would be wary of them.
Interesting. Why? Is it security? The ones I've listed are open source and auditable. I'm confident that they won't steal my API keys. Msty has a lot of advanced functionality that I haven't seen in other interfaces like allowing you to compare responses between different LLMs, export the entire conversation to Markdown, and edit the LLM's response to manage context. It also sidesteps the problem of '[provider] doesn't have a desktop app' because you can use any provider API.
Re: Gemini 3
#574Here are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/
Re: Gemini 3
#575Good. That said, I wonder if those models are still LLMs.
Re: Gemini 3
#576Earlier quoted context omitted.
I was interested (and slightly disappointed) to read that the knowledge cutoff for Gemini 3 is the same as for Gemini 2.5: January 2025. I wonder why they didn't train it on more recent data. Is it possible they use the same base pre-trained model and just fine-tuned and RL-ed it better (which, of course, is where all the secret sauce training magic is these days anyhow)? That would be odd, especially for a major ver…
The model card says: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge... > This model is not a modification or a fine-tune of a prior model. I'm curious why they decided not to update the training data cutoff date too.
Re: Gemini 3
#577Earlier quoted context omitted.
Every time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?
This is a list of questions and answers that was created by different people. The questions AND the answers are public. If the LLM manages through reasoning OR memory to repeat back the answer then they win. The scores represent the % of correct answers they recalled.
You could question how well this works, but it’s not like the answers are just hanging out on the public internet.
Re: Gemini 3
#578Earlier quoted context omitted.
Perception seems to be one of the main constraints on LLMs that not much progress has been made on. Perhaps not surprising, given perception is something evolution has worked on since the inception of life itself. Likely much, much more expensive computationally than it receives credit for.
I strongly suspect it's a tokenization problem. Text and symbols fit nicely in tokens, but having something like a single "dog leg" token is a tough problem to solve.
Re: Gemini 3
#579I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk:
> The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you the bother of washing them yourself; video recorders watched tedious television for you, thus saving you the bother of looking at it yourself. Electric Monks believed things for you, thus saving you what was becoming an increasingly onerous task, that of believing all the things the world expected you to believe.
- from "Dirk Gently's Holistic Detective Agency"
Re: Gemini 3
#580Earlier quoted context omitted.
I think it's fun to see what is not even considered magic anymore today.
People would have had a heart attack if they saw this 5 years ago for the first time. Now artificial brains are “meh” :)
It scares the absolute shit out of everyone.
It's clear far beyond our little tech world to everyone this is going to collapse our entire economic system, destroy everyone's livelihoods, and put even more firmly into control the oligarchic assholes already running everything and turning the world to shit.
I see it in news, commentary, day to day conversation. People get it's for real this time and there's a very real chance it ends in something like the Terminator except far worse.