Live data from Hacker News

Gemini 3

blog.google

641–650 of 1001 posts

Re: Gemini 3

#641
post #487

Here are my notes and pelican benchmark, including a new, harder benchmark because the old one was getting too easy: https://simonwillison.net/2025/Nov/18/gemini-3/

Considering how important this benchmark has become to the judgement of state of the art AI models, I imagine each AI lab has a dedicated 'pelican guy', a a highly accomplished and academically credentialed person, who's working around the clock on training the model to make better and better SVG pelicans on bikes.

That would mean my dastardly scheme has finally come to fruition: https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

Re: Gemini 3

#642
post #134

Earlier quoted context omitted.

These prediction markets are so ripe for abuse it's unbelievable. People need to realize there are real people on the other side of these bets. Brian Armstong, CEO of Coinbase intentionally altered the outcome of a bet by randomly stating "Bitcoin, Ethereum, blockchain, staking, Web3" at the end of an earnings call. These types of bets shouldn't be allowed.

The point of prediction markets isn't to be fair. They are not the stock market. The point of prediction markets is to predict. They provide a monetary incentive for people who are good at predicting stuff. Whether that's due to luck, analysis, insider knowledge, or the ability to influence the result is irrelevant. If you don't want to participate in an unfair market, don't participate in prediction markets.

But what's the point of predicting how many times Elon will say "Trump" on an earnings call (or some random event Kalshi or Polymarket make up)? At least the stock market serves a purpose. People will claim "prediction markets are great for price discovery!" Ok. I'm so glad we found out the chance of Nicki Minaj saying "Bible" during some recent remarks. In case you were wondering, the chance peaked at around 45% and she did not say 'bible'! She passed up a great opportunity to buy the "yes" and make a ton of money!

https://kalshi.com/markets/kxminajmention/nicki-minaj/kxmina...

Re: Gemini 3

#643

What I'd prefer over benchmarks is the answer to a simple question: What useful thing can it demonstrably do that its predecessors couldn't?

Keep the bubble expanding for a few months longer.

Re: Gemini 3

#644
GOOGLE: "We have a new product".

REALITY: It's just 3 existing products rolled into one. One of which isn't even a Google product.

- Microsoft Code

- Gemeni

- Chrome Browser

Re: Gemini 3

#645

I love it that there's a "Read AI-generated summary" button on their post about their new AI. I can only expect that the next step is something like "Have your AI read our AI's auto-generated summary", and so forth until we are all the way at Douglas Adams's Electric Monk: > The Electric Monk was a labour-saving device, like a dishwasher or a video recorder. Dishwashers washed tedious dishes for you, thus saving you…

after outsource developer job, we can outsource all of manager job and leaving CEO with AI agentic code as its servant

Re: Gemini 3

#646

Earlier quoted context omitted.

and models are still pretty bad at playing tic-tac-toe, they can do it, but think way too much it's easy to focus on what they can't do

Everything is about context. When you just ask non-concrete task it's still have to parse your input and figure what is tic-tac-toe in this context and what exactly you expect it to do. This is why all "thinking". Ask it to implement tic-tac-toe in Python for command line. Or even just bring your own tic-tac toe code. Then make it imagine playing against you and it's gonna be fast and reliable.

prompt was very concrete: draw a tic tac toe ASCII table and let's play. gemini 2.5 thought for pages particular moves

Re: Gemini 3

#648
post #326

Earlier quoted context omitted.

Why should it have to be expensive computationally? How do brains do it with such a low amount of energy? I think catching the brain abilities even of a bug might be very hard, but that does not mean that there isn't a way to do it with little computational power. It requires having the correct structures/models/algorithms or whatever is the precise jargon.

This is the million dollar question. I'm not qualified to answer it, and I don't really think anyone out there has the answer yet. My armchair take would be that watt usage probably isn't a good proxy for computational complexity in biological systems. A good piece of evidence for this is from the C. elegans research that has found that the configuration of ions within a neuron--not just the electrical charge on the…

My armchair is equally comfy, and I have an actual paper to point to:

Jaxley: Differentiable simulation enables large-scale training of detailed biophysical models of neural dynamics [1]

They basically created sofware to simulate real neurons and ran some realistic models to replicate typical AI learning tasks:

"The model had nine different channels in the apical and basal dendrite, the soma, and the axon [39], with a total of 19 free parameters, including maximal channel conductances and dynamics of the calcium pumps."

So yeah, real neurons are a bit more complex then ReLU or Sigmoid.

[1] https://www.biorxiv.org/content/10.1101/2024.08.21.608979v2....

Re: Gemini 3

#649

I am personally impressed by the continued improvement in ARC-AGI-2, where Gemini 3 got 31.1% (vs ChatGPT 5.1's 17.6%). To me this is the kind of problem that does not lend itself well to LLMs - many of the puzzles test the kind of thing that humans intuit because of millions of years of evolution, but these concepts do not necessarily appear in written form (or when they do, it's not clear how they connect to specif…

What I would do if I was in the position of a large company in this space is to arrange an internal team to create an ARC replica, covering very similar puzzles and use that as part of the training. Ultimately, most benchmarks can be gamed and their real utility is thus short-lived. But I think this is also fair to use any means to beat it.

> internal team to create an ARC replica, covering very similar puzzles

they can target benchmark directly, not just replica. If google or OAI are bad actors, they already have benchmark data from previous runs.

Re: Gemini 3

#650

Well, I tried a variation of a prompt I was messing with in Flash 2.5 the other day in a thread about AI-coded analog clock faces. Gemini Pro 3 Preview gave me a result far beyond what I saw with Flash 2.5, and got it right in a single shot.[0] I can't say I'm not impressed, even though it's a pretty constrained example. > Please generate an analog clock widget, synchronized to actual system time, with hands that upd…

That is not the same prompt as the other person was using. In particular this doesn't provide the time to set the clock to, which makes the challenge a lot simpler. This also includes javascript.

The prompt the other person was using is:

``` Create HTML/CSS of an analog clock showing ${time}. Include numbers (or numerals) if you wish, and have a CSS animated second hand. Make it responsive and use a white background. Return ONLY the HTML/CSS code with no markdown formatting. ```

Which is much more difficult.

For what it's worth, I supplied the same prompt as the OG clock challenge and it utterly failed, not only generating a terrible clock, but doing so with a fair bit of typescript: https://ai.studio/apps/drive/1c_7C5J5ZBg7VyMWpa175c_3i7NO7ry...

Post reply on HN