Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

11–20 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#11
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Yeah for my agent gemini 2.5 flash performs similar in quality to gpt4.1 and it's way faster and cheaper.

Re: Gemini 3.0 spotted in the wild through A/B testing

#12
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I gave up on Gemini because I couldn't stop the glazing. I don't need to be told what can incredible insight I have made and why my question gets to the heart of the matter every time I ask something.

Re: Gemini 3.0 spotted in the wild through A/B testing

#13
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I tend to find it competitive, but slightly worse on average. But they each have their strengths and weaknesses. I tend to flip between them more than I do search engines.

Re: Gemini 3.0 spotted in the wild through A/B testing

#14
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini is the only model that can provide consistent solution to theoretical physics problems and output it into LaTeX document.

Re: Gemini 3.0 spotted in the wild through A/B testing

#15
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I find Claude and Gemini to be wildly inferior to ChatGPT when it comes to doing searches to establish grounding. Gemini seems to do a handful of searches and then make shit up, where ChatGPT will do dozens or even hundreds of searches - and do searches based on what it finds in earlier ones.

Re: Gemini 3.0 spotted in the wild through A/B testing

#16
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

What application are you using it with? I find this to be very important, for instance it has always SUCKED for me in Copilot (copilot has always kind of sucked for me, but Gemini has managed to regularly completely destroy entire files).

How often do you encounter loops?

Re: Gemini 3.0 spotted in the wild through A/B testing

#17
post #12
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I gave up on Gemini because I couldn't stop the glazing. I don't need to be told what can incredible insight I have made and why my question gets to the heart of the matter every time I ask something.

"Of course! That's an excellent reply to my comment!"

Joking obviously but I've noticed this too, I put up with it because the output is worth it.

Re: Gemini 3.0 spotted in the wild through A/B testing

#18
post #12
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I gave up on Gemini because I couldn't stop the glazing. I don't need to be told what can incredible insight I have made and why my question gets to the heart of the matter every time I ask something.

With AI studio there's a system prompt where you can tell it to stop the sycophancy.

But yeah it does do that otherwise. At one point it told me I'm a genius.

Re: Gemini 3.0 spotted in the wild through A/B testing

#19
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Yes. Jules even writes more testable code, but people I know regularly use codex because it will bang its head against the wall and eventually give you a working implementation even though it took longer.

Re: Gemini 3.0 spotted in the wild through A/B testing

#20
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini was good when the thinking tokens were shown to the user. As soon as Google replaced those with some thought summary, I stopped finding it as useful. Previously, the thoughts were so organized that I would often read those instead of the final answer.
Post reply on HN