Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

81–90 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#81
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I completely disagree. For me the best for bulk coding (with very good instructions) is Sonnet 4.5. Then GPT-5 codex is slower but better guessing what I want with tiny prompts. Gemini 2.5 Pro is good to review large codebases but for real work usually gets confused a lot, not worth it. (even though I was forced to pay for it by Google, I rarely use it).

But the past few days I started getting an "AI Mode" in Google Search that rocks. Way better than GPT-5 or Sonnet 4.5 for figuring out things and planning. And I've been using without my account (weird, but I'm not complaining). Maybe this is Gemini 3.0. I would love for it to be good at coding. I'm near limits on my Anthropic and OpenAI accounts.

Re: Gemini 3.0 spotted in the wild through A/B testing

#82
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I prefer it too, but I find it a bit too wordy. It loves to build narratives. I think this is a common theme with all of Google’s LLMs. Gemma 27B is by far the best in its class for article generation.

Re: Gemini 3.0 spotted in the wild through A/B testing

#83
post #76

Earlier quoted context omitted.

Looking at the responses. How the F have people so wildly different opinions on the relative performance of the same systems?

Different prompts/approaches? I "grew up", as it were, on StackOverflow, when I was in my early dev days and didn't have a clue what I was doing I asked question after question on SO and learned very quickly the difference between asking a good question vs asking a bad one There is a great Jon Skeet blog post from back in the day called "Writing the perfect question" - https://codeblog.jonskeet.uk/2010/08/29/writing-…

Sure but if one is bad at asking questions they would be consistently bad across chatbots

Re: Gemini 3.0 spotted in the wild through A/B testing

#84
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I had the same feeling when 2.5 pro was initially released, but it seemed like after a while they quantized the model.

Re: Gemini 3.0 spotted in the wild through A/B testing

#85

Earlier quoted context omitted.

> consistently found Gemini to be better than ChatGPT, Claude and Deepseek I used Pro Mode in ChatGPT since it was available, and tried Claude, Gemini, Deepseek and more from time to time, but none of them ever get close to Pro Mode, it's just insanely better than everything. So when I hear people comparing "X to ChatGPT", are you testing against the best ChatGPT has to offer, or are you comparing it to "Auto" and ca…

It seems you also did not compare ChatGPT to the best offers of the competitors, as you did not mention Gemini Deepthink mode which is Google's alternative to GPT's Pro mode.

TBH, I always forget that Deepthink is even an option. It's powerful, but not exactly conspicuous.

Re: Gemini 3.0 spotted in the wild through A/B testing

#86
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

> consistently found Gemini to be better than ChatGPT, Claude and Deepseek I used Pro Mode in ChatGPT since it was available, and tried Claude, Gemini, Deepseek and more from time to time, but none of them ever get close to Pro Mode, it's just insanely better than everything. So when I hear people comparing "X to ChatGPT", are you testing against the best ChatGPT has to offer, or are you comparing it to "Auto" and ca…

Yeah, ChatGPT “auto”, at least when it ends up routing to gpt-5-chat, is a slopfest. I discounted gpt-5 early on due to that experience.

Now I have my model selector permanently on “Thinking”. (I don’t even know what type of questions I’d ask the non-thinking one.)

Re: Gemini 3.0 spotted in the wild through A/B testing

#87
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

We extensively benchmark frontier models at $DAYJOB and Gemini 2.5 is the uncontested king outside of a few narrow use cases. Tracks with the rumor that Google has the best pretraining and falls short only in tuning/alignment. Eagerly anticipating Gemini 3 as 2.5, while king of the hill, still has lots of room for improvement!

Edit: narrow use cases are roughly "true reasoning" (GPT-5) and Python script writing (the Claudes)

Re: Gemini 3.0 spotted in the wild through A/B testing

#89
There are a lot more of these Gemini 3 examples out on twitter right now.

After seeing them, I bought Google stock. What shocks me about its output is it actually feels like it's producing net new creative designs, not just regurgitated template output. Its extremely hard to design in code in a way that produces consistent, beautiful output, but it seems to be achieving it.

That combined with Google being the only one in the core model space that is fully vertically integrated with their own hardware makes me feel extremely bullish on their success in the AI race.

Re: Gemini 3.0 spotted in the wild through A/B testing

#90
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I am curious what your background is. I also almost exclusively use Gemini 2.5, and my PhD colleagues in comp sci do the same. However it seems like the general public, or people outside this bubble are more likely to use ChatGPT or Claude.

I wonder if it has something to do with the level of abstraction and questions that you give to Gemini, which might be related to the profession or way of typing.

Post reply on HN