Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

31–40 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#31
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Looking at the responses. How the F have people so wildly different opinions on the relative performance of the same systems?

Re: Gemini 3.0 spotted in the wild through A/B testing

#32
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I find Gemini excels at greenfield, big picture tasks. I use Sonnet and Codex for implementation.

Re: Gemini 3.0 spotted in the wild through A/B testing

#33
post #9
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

My pet theory is that Gemini's training is, more than others, focused on rewriting and pulling out facts from data. (As well as being cheap to run). Since the biggest use is the Google AI generated search results

It doesn't perform nearly as well as Claude or even Codex for my programming tasks though

Re: Gemini 3.0 spotted in the wild through A/B testing

#34
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Agreed. There seems to be some very strong anti-Google force on HN. I guess there's just a lot of astroturfing in this area.

Re: Gemini 3.0 spotted in the wild through A/B testing

#35
post #26
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I swear HN commenters say this about every frontier model.

[dead]

Re: Gemini 3.0 spotted in the wild through A/B testing

#39
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

We've moved to it for our clinical workflow agents. Great quality, better pricing and performance compared to Anthropic.

Re: Gemini 3.0 spotted in the wild through A/B testing

#40
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Depends on the task, our tastes, and our workflow. In my case:

For writing and editorial work, I use Gemini 2.5 Pro (Sonnet seems simply worse, while GPT5 too opinionated).

For coding, Sonnet 4.5 (usually).

For brainstorming and background checks, GPT5 via ChatGPT.

For data extraction, GPT5. (Seems to be the best at this "needle in a haystack".)

Post reply on HN