Live data from Hacker News

Gemini 3

blog.google

801–810 of 1001 posts

Re: Gemini 3

#801
post #698
post #690

Hoping someone here may know the answer to this, but do any of the benchmarks that exist currently account for false answers in any meaningful way, other than it would in a typical test (ie, if I give any answer at all it is better than saying "I don't know" as the answer I give at least has a chance of being correct(which in the real world is bad))? I want an LLM that tells me when it doesn't know something. If it g…

Those numbers are too good to expect. If 90% right 10% wrong is the baseline would you take as an improvement: - 80% right 18% I don't know 2% wrong - 50%/48%/2% - 10%/90%/0% - 80%/15%/5% The general point being that to reduce wrong answers you will need to accept some reduction in right answers if you want the change to only be made through trade-offs. Otherwise you just say "I'd like a better system" and that is ra…

I think you may have misread. They stated that they'd be willing to go from 90% correct to 10% correct for this tradeoff.

Re: Gemini 3

#802
post #435

Earlier quoted context omitted.

The requested prompt does not exist or you do not have access. If you believe the request is correct, make sure you have first allowed AI Studio access to your Google Drive, and then ask the owner to share the prompt with you.

I thought this was a joke at first. It actually needs drive access to run someone else's prompt. Wild.

After ChatGPT accidentally indexed everyones shared chats (and had a cache collision in their chat history early on) and Meta build a UI flow that filled a public feed full of super private chats... seems like a good move to use a battle tested permission system.

Re: Gemini 3

#803
post #745

Earlier quoted context omitted.

So when does the developer admit defeat? Do we have a benchmark for that yet?

According to a bunch of philosophers ( https://ai-2027.com/ ), doom is likely imminent. Kokotajlo was on Breaking Points today. Breaking Points is usually less gullible, but the top comment shows that "AI" hype strategy detection is now mainstream ( https://www.youtube.com/watch?v=zRlIFn0ZIlU ): AI researcher: "Just another trillion dollars. This time we'll reach superintelligence, I swear."

Every Ai researcher calls it quits one YOLO run away from inventing a machine that turns all matter in the Universe into paperclips

Re: Gemini 3

#804

Earlier quoted context omitted.

[flagged]

[flagged]

1) Models learn these patterns from common human usage. They are in the wild, and as such there will be people who use them naturally.

2) Now, given its for-some-reason-ubiquitous choice by models, it is also a phrasing that many more people are exposed to, every day.

Language is contagious. This phrasing is approaching herd levels, meaning models trained from up-to-the-moment web content will start to see it as less distinctly salient. Eventually, there will be some other high-signal novel phrase with high salience, and the attention heads will latch on to it from the surrounding context, and then that will be the new AI shibboleth.

It's just how language works. We see it in the mixes between generations when our kids pick up new lingo, and then it stops being in-group for them when it spreads too far.. Skibidi, 6 7, etc.

It's just how language works, and a generation ago the internet put it on steroids. Now? Even faster.

Re: Gemini 3

#805
The most devastating news out of this announcement is that Vending-Bench 2 came out and it has significantly less clanker[0] meltdowns than the first one. I mean, seriously? Not even one run where the model tried to stock goods that hadn't arrived yet, only for it to eventually try and fail to shut down the business, and then e-mail the FBI about the $2 daily fee being deducted from the bot?

[0] Fake racial slur for a robot, LLM chatbot, or other automated system

Re: Gemini 3

#806

Earlier quoted context omitted.

[flagged]

I usually ask a simple question that ALL the models get wrong: List of mayor of my city [Londrina]. ALL the models (offine) get wrong. And I mean, all the models. The best that I could, it's o3 I believe, saying it couldn't give a good answer for that, and told to access the city website. Gemini 3 somehow is able to give a list of mayors, including details on who got impeached, etc. This should be a simple answer, be…

thanks for sharing, very interesting example

Re: Gemini 3

#807
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

[flagged]

It tells me that the benchmark is probably leaking into training data, and going to the benchmark site :

> Model was published after the competition date, making contamination possible.

Aside from eval on most of these benchmarks being stupid most of the time, these guys have every incentive to cheat - these aren't some academic AI labs, they have to justify hundreds of billions being spent/allocated in the market.

Actually trying the model on a few of my daily tasks and reading the reasoning traces all I'm seeing is same old, same old - Claude is still better at "getting" the problem.

Re: Gemini 3

#808

Earlier quoted context omitted.

[flagged]

I usually ask a simple question that ALL the models get wrong: List of mayor of my city [Londrina]. ALL the models (offine) get wrong. And I mean, all the models. The best that I could, it's o3 I believe, saying it couldn't give a good answer for that, and told to access the city website. Gemini 3 somehow is able to give a list of mayors, including details on who got impeached, etc. This should be a simple answer, be…

Ha, I just did the same with my hometown (Guaiba, RS), a city that is 1/6th of Londrina, and its wikipedia page in English hasn't been updated in years, and still has the wrong mayor (!).

Gemini 3 nailed on the first try, included political affiliation, and added some context on who they competed with and won over in each of the last 3 elections. And I just did a fun application with AI Studio, and it worked on first shot. Pretty impressive.

(disclaimer: Googler, but no affiliation with Gemini team)

Re: Gemini 3

#809

Earlier quoted context omitted.

[flagged]

I usually ask a simple question that ALL the models get wrong: List of mayor of my city [Londrina]. ALL the models (offine) get wrong. And I mean, all the models. The best that I could, it's o3 I believe, saying it couldn't give a good answer for that, and told to access the city website. Gemini 3 somehow is able to give a list of mayors, including details on who got impeached, etc. This should be a simple answer, be…

Pure fact-based, niche questions like that aren't really the focus of most providers any more from what I've heard, since they can be solved more reliably by integrating search tools (and all providers now have search).

I wouldn't be surprised if the smallest models can answer fewer such (fact-only) questions over time offline as they distill/focus them more thoroughly on logic etc.

Re: Gemini 3

#810
post #443

Earlier quoted context omitted.

After few more attempts longer animation with a story from my gamedev inspired mind: https://codepen.io/Runway/pen/zxqzPyQ PS: but yeah thats attempt #20 or something.

Wow looks like total shit and eventually very hard to take on and actually improve it, given the convoluted code it generated, YET people are impressed. What world are we living in...

You are missing the point of this exercise. This is not about code quality - its about capacity of model to generate visuals with no guidance.

For the code quality it can really be as good or as bad ad as you desire. In this case it is what it is because I put zero effort into it.

Post reply on HN