Live data from Hacker News

Gemini 3

blog.google

831–840 of 1001 posts

Re: Gemini 3

#831
Created a summary of comments from this thread about 15 hours after it had been posted and had 814 comments with gemini-3-pro and gpt-5.1 using this script [1]:

- gemini-3-pro summary: https://gist.github.com/primaprashant/948c5b0f89f1d5bc919f90...

- gpt-5.1 summary: https://gist.github.com/primaprashant/3786f3833043d8dcccae4b...

Summary from GPT 5.1 is significantly longer and more verbose compared to Gemini 3 Pro (13,129 output tokens vs 3,776). Gemini 3 summary seems more readable, however, GPT 5.1 one has interesting insights missed by Gemini.

Last time I did this comparison at the time of GPT 5 release [2], the summary from Gemini 2.5 Pro was way better and readable than the GPT 5 one. This time the readability of Gemini 3 summary still seems great while GPT 5.1 feels a bit more improved but not there quite yet.

[1]: https://gist.github.com/primaprashant/f181ed685ae563fd06c49d...

[2]: https://news.ycombinator.com/item?id=44835029

Re: Gemini 3

#832
post #443

Earlier quoted context omitted.

After few more attempts longer animation with a story from my gamedev inspired mind: https://codepen.io/Runway/pen/zxqzPyQ PS: but yeah thats attempt #20 or something.

Wow looks like total shit and eventually very hard to take on and actually improve it, given the convoluted code it generated, YET people are impressed. What world are we living in...

You can criticize the code but "wow looks like total shit" is such an embarrassing thing to say considering the context. Imagine going back a few years and show them a tool outputting this from text. No-one would believe it.

Re: Gemini 3

#833

Earlier quoted context omitted.

I also used Gemini 3 Pro Preview. It finished it 271s = 4m31s. Sadly, the answer was wrong. It also returned 8 "sources", like stackexchange.com, youtube.com, mpmath.org, ncert.nic.in, and kangaroo.org.pk, even though I specifically told it not to use websearch. Still a useful tool though. It definitely gets the majority of the insights. Prompt: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Why is this sad. You should bw rooting for these LLMs to be as bad as possible..

> You should bw rooting for these LLMs to be as bad as possible..

Why?

Re: Gemini 3

#834

Earlier quoted context omitted.

Why is this sad. You should bw rooting for these LLMs to be as bad as possible..

Generally, any expert hopes their tool/paintbrush/etc is as performant as possible.

And in general I'm all for increasing productivity, in all areas of the economy.

Re: Gemini 3

#835

Earlier quoted context omitted.

How is it useful other than for people making money off token outout. Continue to fry your brain.

They’re fantastic learning tools, for a start. What you get out of them is proportional to what you put in. You’ve probably heard of the Luddites, the group who destroyed textile mills in the early 1800s. If not: https://en.wikipedia.org/wiki/Luddite Luddites often get a bad rap, probably in large part because of employer propaganda and influence over the writing of history, as well as the common tendency of people t…

Arguably the Luddites don't get a bad enough rep. The lump of labour fallacy was as bad then as it is now or at any other time.

https://en.wikipedia.org/wiki/Lump_of_labour_fallacy

Re: Gemini 3

#836
post #65

My favorite benchmark is to analyze a very long audio file recording of a management meeting and produce very good notes along with a transcript labeling all the speakers. 2.5 was decently good at generating the summary, but it was terrible at labeling speakers. 3.0 has so far absolutely nailed speaker labeling.

I'd do the transcript and the summary parts separately. Dedicated audio models from vendors like ElevenLabs or Soniox use speaker detection models to produce an accurate speaker based transcript while I'm not necessarily sure that Google's models do so, maybe they just hallucinate the speakers instead.

Agreed. I don’t see the need for Gemini to be able to do this task, although it should be able to offload it to another model.

Re: Gemini 3

#837
post #293

Out of curiosity, I gave it the latest project euler problem published on 11/16/2025, very likely out of the training data Gemini thought for 5m10s before giving me a python snippet that produced the correct answer. The leaderboard says that the 3 fastest human to solve this problem took 14min, 20min and 1h14min respectively Even thought I expect this sort of problem to very much be in the distribution of what the mo…

I also used Gemini 3 Pro Preview. It finished it 271s = 4m31s. Sadly, the answer was wrong. It also returned 8 "sources", like stackexchange.com, youtube.com, mpmath.org, ncert.nic.in, and kangaroo.org.pk, even though I specifically told it not to use websearch. Still a useful tool though. It definitely gets the majority of the insights. Prompt: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%...

Terrence Tao claims [0] contributions by the public are counter-productive since the energy required to check a contribution outweighs its benefit:

> (for) most research projects, it would not help to have input from the general public. In fact, it would just be time-consuming, because error checking

Since frontier LLMs make clumsy mistakes, they may fall into this category of 'error-prone' mathematician whose net contributions are actually negative, despite being impressive some of the time.

[0] https://www.youtube.com/watch?v=HUkBz-cdB-k&t=2h59m33s

Re: Gemini 3

#838

DeepMind page: https://deepmind.google/models/gemini/ Gemini 3 Pro DeepMind Page: https://deepmind.google/models/gemini/pro/ Developer blog: https://blog.google/technology/developers/gemini-3-developer... Gemini 3 Docs: https://ai.google.dev/gemini-api/docs/gemini-3 Google Antigravity: https://antigravity.google/

Also recently: Code Wiki: https://codewiki.google/

Re: Gemini 3

#839

I'm sure this is a very impressive model, but gemini-3-pro-preview is failing spectacularly at my fairly basic python benchmark. In fact, gemini-2.5-pro gets a lot closer (but is still wrong). For reference: gpt-5.1-thinking passes, gpt-5.1-instant fails, gpt-5-thinking fails, gpt-5-instant fails, sonnet-4.5 passes, opus-4.1 passes (lesser claude models fail). This is a reminder that benchmarks are meaningless – you…

> This is a reminder that benchmarks are meaningless – you should always curate your own out-of-sample benchmarks. Yeah I have my own set of tests and the results are a bit unsettling in the sense that sometimes older models outperform newer ones. Moreover, they change even if officially the model doesn't change. This is especially true of Gemini 2.5 pro that was performing much better on the same tests several month…

I wonder whether it could be related to some kind of over-fitting, i.e. a prompting style that tends to work better with the older models, but performs worse with the newer ones.

Re: Gemini 3

#840
post #711

I just gave it a short description of a small game I had an idea for. It was 7 sentences. It pretty much nailed a working prototype, using React, clean css, Typescript and state management. It event implemented a Gemini query using the API for strategic analysis given a game state. I'm more than impressed, I'm terrified. Seriously thinking of a career change.

I find it funny to find this almost exact same post in every new model release thread. Yet here we are - spending the same amount of time, if not more, finishing the rest of the owl.

[deleted]
Post reply on HN