Live data from Hacker News

Gemini 3

blog.google

871–880 of 1001 posts

Re: Gemini 3

#871
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

definitely uses a lot of tooling. From "thinking":

> I'm now writing a Python script to automate the summation computation. I'm implementing a prime sieve and focusing on functions for Rm and Km calculation [...]

Re: Gemini 3

#872
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

Does it matter if it is out of the training data? The models integrate web search quite well. What if they have an internal corpus of new and curated knowledge that is constantly updated by humans and accessed in a similar manner? It could be active even if web search is turned off. They would surely add the latest Euler problems with solutions in order to show off in benchmarks.

you can disable search.

just create a different problem if you don't believe it.

Re: Gemini 3

#873

Earlier quoted context omitted.

they are not meaningless, but when you work a lot with LLMs and know them VERY well, then a few varied, complex prompts tell you all you need to know about things like EQ, sycophancy, and creative writing. I like to compare them using chathub using the same prompts Gemini still calls me "the architect" in half of the prompts. It's very cringe.

Gemini still calls me "the architect" in half of the prompts. It's very cringe. Can't say I've ever seen this in my own chats. Maybe it's something about your writing style?

it absolutely does. and human employees don't call me "the architect." that's the point.

Re: Gemini 3

#874
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

[flagged]

> To succeed this well in math, you can't just do better probabilistic generation, you need verifiable search.

You say "probabilistic generation" like it's some kind of a limitation. What is exactly the limiting factor here? [(0.9999, "4"), (0.00001, "four"), ...] is a valid probability distribution. The sampler can be set to always choose "4" in such cases.

Re: Gemini 3

#875

Earlier quoted context omitted.

What's the benchmark?

I don't think it would be a good idea to publish it on a prime source of training data.

but they've asked all the AI models this question. Whatever you tell an AI model is also in its training data

Re: Gemini 3

#876

Earlier quoted context omitted.

I've been working with it, and so far it's been very impressive. Better than Opus in my feels, but I have to test more, it's super early days

What I usually try to test with is try to get them do full scalable SaaS application from scratch... It seemed very impressive in how it did the early code organization using Antigravity, but then at some point, all of sudden it started really getting stuck and constantly stopped producing and I had to trigger continue, or babysit it. I don't know if I could've been doing something better, but that was just my experi…

It sounds like an API issue more than anything. I was working with it through cursor on a side project, and it did better than all previous models at following instructions, refactoring, and UI-wise it has some crazy skills.

What really impressed me was when I told it that I wanted a particular component’s UI to be cleaned up but I didn’t know how exactly, just wanted to use its deep design expertise to figure it out, and it came up with a UX that I would’ve never thought of and that was amazing.

Another important point is that the error rate for my session yesterday was significantly lower than when I’ve used any other model.

Today I will see how it does when I use it at work, where we have a massive codebase that has particular coding conventions. Curious how it does there.

Re: Gemini 3

#877
post #837

Earlier quoted context omitted.

I also used Gemini 3 Pro Preview. It finished it 271s = 4m31s. Sadly, the answer was wrong. It also returned 8 "sources", like stackexchange.com, youtube.com, mpmath.org, ncert.nic.in, and kangaroo.org.pk, even though I specifically told it not to use websearch. Still a useful tool though. It definitely gets the majority of the insights. Prompt: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Terrence Tao claims [0] contributions by the public are counter -productive since the energy required to check a contribution outweighs its benefit: > (for) most research projects, it would not help to have input from the general public. In fact, it would just be time-consuming, because error checking Since frontier LLMs make clumsy mistakes, they may fall into this category of 'error-prone' mathematician whose net c…

Unlike general public the models can be trained. I mean if you train a member of general public, you've got a specialist, who is no longer a member of general public.

Re: Gemini 3

#878

I asked Gemini to write "a comment response to this thread. I want to start an intense discussion". Gemini 3: The cognitive dissonance in this thread is staggering. We are sitting here cheering for a model that effectively closes the loop on Google’s total information dominance, while simultaneously training our own replacements. Two things in this thread should be terrifying, yet are being glossed over in favor of "…

> We are cheering for a product sold back to us at a 60% markup (input costs up to $2.00/M) that was built on our own private correspondence. That feels like something between a hallucination and an intentional fallacy that popped up because you specifically said "intense discussion". The increase is 60% on input tokens from the old model, but it's not a markup, and especially not "sold back to us at X markup". I've…

60% probably felt like a lot to Gemini. However, I liked the doomerism and how google was using our data to train its models.

Nonetheless, Gemini 3 failed this test. It failed to start a discussion. Its points were shallow, and too aiesque.

Re: Gemini 3

#879
post #366

Static Pelican is boring. First attempt: Generate SVG animation of following: 1 - There is High fantasy mage tower with a top window a dome 2 - Green goblin come in front of tower with a torch 3 - Grumpy old mage with beard appear in a tower window in high purple hat 4 - Mage sends fireball that burns goblin and all screen is covered in fire. Camera view must be from behind of goblin back so we basically look at towe…

we are returning to flash animations after 20 years

Nature is healing!

But seriously, we lost a lot when Flash was killed. It was an era of accessible animation and games like Newgrounds and Homestar Runner, that had no ready replacement.

Re: Gemini 3

#880

Earlier quoted context omitted.

According to a bunch of philosophers ( https://ai-2027.com/ ), doom is likely imminent. Kokotajlo was on Breaking Points today. Breaking Points is usually less gullible, but the top comment shows that "AI" hype strategy detection is now mainstream ( https://www.youtube.com/watch?v=zRlIFn0ZIlU ): AI researcher: "Just another trillion dollars. This time we'll reach superintelligence, I swear."

Every Ai researcher calls it quits one YOLO run away from inventing a machine that turns all matter in the Universe into paperclips

Do NOT follow this link:

https://www.decisionproblem.com/paperclips/

Post reply on HN