Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

41–50 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#41
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

> consistently found Gemini to be better than ChatGPT, Claude and Deepseek I used Pro Mode in ChatGPT since it was available, and tried Claude, Gemini, Deepseek and more from time to time, but none of them ever get close to Pro Mode, it's just insanely better than everything. So when I hear people comparing "X to ChatGPT", are you testing against the best ChatGPT has to offer, or are you comparing it to "Auto" and ca…

It seems you also did not compare ChatGPT to the best offers of the competitors, as you did not mention Gemini Deepthink mode which is Google's alternative to GPT's Pro mode.

Re: Gemini 3.0 spotted in the wild through A/B testing

#42
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini is theoretically better, but I find it's very unsteerable. Combine that with the fact it struggles with tool use and character-level issues - and it can be challenging to use despite being "smarter".

Re: Gemini 3.0 spotted in the wild through A/B testing

#43
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I use GPro 2.5 exclusively for coding anything difficult, and Claude Opus otherwise.

Between the two, 100% of my code is written by AI now, and has been since early July. Total gamechanger vs. earlier models, which weren't usable for the kind of code I write at all.

I do NOT use either as an "agent." I don't vibe code. (I've tried Claude Code, but it was terrible compared to what I get out of GPro 2.5.)

Re: Gemini 3.0 spotted in the wild through A/B testing

#44
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini was good when the thinking tokens were shown to the user. As soon as Google replaced those with some thought summary, I stopped finding it as useful. Previously, the thoughts were so organized that I would often read those instead of the final answer.

These were extremely helpful to read for insights on how to go back and retry different prompts instead, IMHO. I find it to be a significant step back in usability to lose those although I can understand the argument that they weren't directly useful on their own outside of that use case.

Re: Gemini 3.0 spotted in the wild through A/B testing

#45
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I used Gemini at work, and would probably agree with your sentiment. For personal usage though, I've stuck with ChatGPT (pro subscriber).. the ChatGPT app has become my default 'ask a question' versus google, and I never reach for Gemini in personal time.

Re: Gemini 3.0 spotted in the wild through A/B testing

#46

Earlier quoted context omitted.

Yes. Jules even writes more testable code, but people I know regularly use codex because it will bang its head against the wall and eventually give you a working implementation even though it took longer.

Maybe because Jules is made by Google and 95% of Google products end up dead as soon as the product manager gets a promotion?

Watch them retire Jules as part of Gemini 3.0 release.

Re: Gemini 3.0 spotted in the wild through A/B testing

#47
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini is theoretically better, but I find it's very unsteerable. Combine that with the fact it struggles with tool use and character-level issues - and it can be challenging to use despite being "smarter".

I agree with the steerable angle, it's like driving a fast car with no traction control

However if you get the hang of it, it can be very powerful

Re: Gemini 3.0 spotted in the wild through A/B testing

#48
This is super exciting. Gemini 2.5 pro was starting to feel like it's lagging behind a little bit; or at least it's still near the best but 3.0 had to be coming along.

It's my goto coder; it just jives better with me than claude or gpt. Better than my home hardware can handle.

What I really hope for 3.0. Their context length is real 1 million. In my experience 256k is the real limit.

Re: Gemini 3.0 spotted in the wild through A/B testing

#49
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini is theoretically better, but I find it's very unsteerable. Combine that with the fact it struggles with tool use and character-level issues - and it can be challenging to use despite being "smarter".

What does it mean for one model to be theoretically better than another?

Re: Gemini 3.0 spotted in the wild through A/B testing

#50
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

gemini used to be the top for me until gpt-5 (web dev with html/js/css + python) ... and also with gpt-5 around it's doing its job, but it's really slow.
Post reply on HN