Live data from Hacker News

Gemini 3

blog.google

671–680 of 1001 posts

Re: Gemini 3

#671

Well, I tried a variation of a prompt I was messing with in Flash 2.5 the other day in a thread about AI-coded analog clock faces. Gemini Pro 3 Preview gave me a result far beyond what I saw with Flash 2.5, and got it right in a single shot.[0] I can't say I'm not impressed, even though it's a pretty constrained example. > Please generate an analog clock widget, synchronized to actual system time, with hands that upd…

"Allow access to Google Drive to load this Prompt." .... why? For what possible reason? No, I'm not going to give access to my privately stored file share in order to view a prompt someone has shared. Come on, Google.

Because most likely (at least according to Hanlon's razor) they somehow decided that using Google Drive as the only persistent storage backing AI studio was a reasonable UX decision.

It probably makes some sense internally in big tech corporation logic (no new data storage agreements on top of the ones the user has already agreed to when signing up for Drive etc.), but as a user, I find it incredibly strange too – especially since the text chats are in some proprietary format I can't easily open on my local GDrive replica, but the images generated or uploaded just look like regular JPEGs and PNGs.

Re: Gemini 3

#672
post #450

Earlier quoted context omitted.

I’d love if anyone could provide examples of such AND(“ground truth”, “absolutely ridiculous”) solutions! Even if they took clever humans a long time to create. I’m curious to explore such fun programming code. But I’m also curious to explore what knowledgeable humans consider to be both “ground truth” as well as “absolutely ridiculous” to create within the usual time constraints.

I'm not explaining myself right. Stockfish is a superhuman chess program. It's routinely used in chess analysis as "ground truth": if Stockfish says you've made a mistake, it's almost certain you did in fact make a mistake[0]. Also, because it's incomparably stronger than even the very best humans, sometimes the moves it suggests are extremely counterintuitive and it would be unrealistic to expect a human to find the…

I would love to examine Stockfish play that seemed extremely counterintuitive but which ended up winning. How can I do so? (I don't inhabit any of the current chess spaces so have no idea where to look, but my son is approaching the age where I can start to teach him...).

That said, chess is such a great human invention. (Go is up there too. And texas no-limit hold'em poker. Those are my top 3 votes for "best human tabletop games ever invented". They're also, perhaps not uncoincidentally, the hardest for computers to be good at. Or, were.)

Re: Gemini 3

#673
post #65

My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.

Parakeet TDT v3 would be really good at that

Yes, this is the best solution for that goal. Use the MacWhisper app + Parakeet 3.

Re: Gemini 3

#674

I am personally impressed by the continued improvement in ARC-AGI-2, where Gemini 3 got 31.1% (vs ChatGPT 5.1's 17.6%). To me this is the kind of problem that does not lend itself well to LLMs - many of the puzzles test the kind of thing that humans intuit because of millions of years of evolution, but these concepts do not necessarily appear in written form (or when they do, it's not clear how they connect to specif…

There's a good chance Gemini 3 was trained on ARG-AGI problems, unless they state otherwise.

ARC-AGI has a hidden private test suite, right ? No model will have access to that set.

Re: Gemini 3

#675
post #574

Earlier quoted context omitted.

Considering how many other "pelican riding a bicycle" comments there are in this thread, it would be surprising if this was not already incorporated in the training data. If not now, soon.

I don't think the big labs would waste their time on it. If a model is great at making the pelican but sucks at all other svg it becomes obvious. But so far the good pelicans are strong indicators of good general SVG ability. Unless training on the pelican increases all SVG ability, then good job.

I absolutely think they would given the amount of money and hype being pumped into it.

Re: Gemini 3

#676

Is there a way to use this without being in the whole google ecosystem? Just make a new account or something?

You could probably do a new account. I have the odd junk google account.

Re: Gemini 3

#677

Earlier quoted context omitted.

after outsource developer job, we can outsource all of manager job and leaving CEO with AI agentic code as its servant

Not sure what you mean here, but the only real jobs at risk from AI right now are middle/upper management. Not a single engineer has ever been laid off because of AI. Any company claiming this is the case is trying to cover up bad decisions. "Were automating with AI" sounds better to investors than "We over hired and now need to downsize" or "We made some bad market bets, now need to free up cash flow"

> Not sure what you mean here, but the only real jobs at risk from AI right now are middle/upper management.

> Not a single engineer has ever been laid off because of AI. Any company claiming this is the case is trying to cover up bad decisions.

I don't suppose these assertions are based on anything. If "AI" reduces the amount of time an engineer spends writing crud, boilerplate, test cases, random scripts, etc., and they have 5% more time to do other things, then all else being equal a project can be done with 5% fewer engineers.

Does AI result in greater productivity for engineers, and does greater productivity per person mean demand can be satisfied with fewer people?

Re: Gemini 3

#678
post #593

Earlier quoted context omitted.

after taking a walk for a bit i decided you’re right. I came to the wrong conclusion. Gemini 3 is incredibly powerful in some other stuff I’ve run. This probably means my test is a little too niche. The fact that it didn’t pass one of my tests doesn’t speak to the broader intelligence of the model per se. While i still believe in the importance of a personalized suite of benchmarks, my python one needs to be down wei…

> This probably means my test is a little too niche. > my python one needs to be down weighted or supplanted. To me, this just proves your original statement. You can't know if an AI can do your specific task based on benchmarks. They are relatively meaningless. You must just try. I have AI fail spectacularly, often, because I'm in a niche field. To me, in the context of AI, "niche" is "most of the code for this is p…

I feel similarly. If you're working with some relatively niche APIs on services that don't get seen by the public, the AI isn't one-shotting anything. But I still find it helpful to generate some crap that I can then feel good about fixing.

Re: Gemini 3

#679
post #364

I have "unlimited" access to both Gemini 2.5 Pro and Claude 4.5 Sonnet through work. From my experience, both are capable and can solve nearly all the same complex programming requests, but time and time again Gemini spits out reams and reams of code so over engineered, that totally works, but I would never want to have to interact with. When looking at the code, you can't tell why it looks "gross", but then you ask…

    but I would never want to have to interact with
That is its job security ;)

Re: Gemini 3

#680

Earlier quoted context omitted.

It says Gemini App, not AI Overviews, AI Mode, etc

They claim AI overviews as having "2 billion users" in the sentences prior. They are clearly trying as hard as possible to show the "best" numbers.

> They are clearly trying as hard as possible to show the "best" numbers.

This isnt a hottake at all. Marketing (iPhone keynotes, product launches) are about showing impressive numbers. It isnt a gotcha you think it is.

Post reply on HN