I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…
Gemini 3.0 spotted in the wild through A/B testing
141–150 of 280 posts
Re: Gemini 3.0 spotted in the wild through A/B testing
#142The sentiment in this thread surprises me a great deal. For me, Gemini 2.5 Pro is markedly worse than GPT-5 Thinking along every axis of hallucinations, rigidity in its self-assured correctness and sycophancy. Claude Opus used to be marginally better but now Claude Sonnet 4.5 is far better, although not quite on par with GPT-5 Thinking. I frequently ask the same question side-by-side to all 3 and the only situation i…
Re: Gemini 3.0 spotted in the wild through A/B testing
#143The sentiment in this thread surprises me a great deal. For me, Gemini 2.5 Pro is markedly worse than GPT-5 Thinking along every axis of hallucinations, rigidity in its self-assured correctness and sycophancy. Claude Opus used to be marginally better but now Claude Sonnet 4.5 is far better, although not quite on par with GPT-5 Thinking. I frequently ask the same question side-by-side to all 3 and the only situation i…
Re: Gemini 3.0 spotted in the wild through A/B testing
#144Earlier quoted context omitted.
Just imagine you’re trying to build a custom D&D campaign for your friends. You might have a fun idea don’t have the time or skills to write yourself that you can have an LLM help out with. Or at least make a first draft you can run with. What do your friends care if you wrote it yourself or used an LLM? The quality bar is going to be fairly low either way, and if it provides some variation from the typical story boo…
LLMs have issues with creative tasks that might not be obvious for light users. Using them for an RPG campaign could work if the bar is low and it's the first couple of times you use it. But after a while, you start to identify repeated patterns and guard rails. The weights of the models are static. It's always predicting what the best association is between the input prompt and whatever tokens its spitting out with…
Re: Gemini 3.0 spotted in the wild through A/B testing
#145I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
I like Gemini 2.5 as a chatbot, but it has been mostly useless as an agent comparing to Claude Code (at least for my complex tasks)
You have to convince it of basic things it refuses to do - no actually you CAN read files outside of the project- try it.
And it'll frequently write \n instead of actually doing a newline when writing files.
It'll straight up ignore/forget a pattern it was JUST properly doing.
Etc.
Re: Gemini 3.0 spotted in the wild through A/B testing
#146Earlier quoted context omitted.
I find Claude and Gemini to be wildly inferior to ChatGPT when it comes to doing searches to establish grounding. Gemini seems to do a handful of searches and then make shit up, where ChatGPT will do dozens or even hundreds of searches - and do searches based on what it finds in earlier ones.
Try "AI Mode" on Google.com (Disclaimer, I recently joined the team that makes this product). It isn't Gemini (the product, those are different orgs) though there may (deliberately left ambiguous) be overlap in LLM level bytes. My recommendation for you in this use-case comes from the fact that AI Mode is a product that is built to be a good search engine first, presented to you in the interface of an AI Chatbot. Rat…
I take no sides; not a fanboy. Only used free Claude and free Gemini Pro 2.5. But some months ago I scoffed at the expression "try it in Google AI Studio" -- that by itself is a branding / marketing failure.
Something like the existing https://ai.google website and with links to the different offerings indeed goes a LONG way. I like that website though it can be done better.
But anyway. Please tell somebody higher up that they are acting like 50 mini companies forced into a single big entity. Google should be better than that.
FWIW, I like Gemini Pro 2.5 best even though I had the free Claude run circles around it sometimes. It one-shot puzzling problems with minimal context multiple times while Gemini was still offering me ideas about how my computer might be malfunctioning if the thing it just hallucinated was not working. Still, most of the time it performs really great.
Re: Gemini 3.0 spotted in the wild through A/B testing
#147Re: Gemini 3.0 spotted in the wild through A/B testing
#148Earlier quoted context omitted.
I find Claude and Gemini to be wildly inferior to ChatGPT when it comes to doing searches to establish grounding. Gemini seems to do a handful of searches and then make shit up, where ChatGPT will do dozens or even hundreds of searches - and do searches based on what it finds in earlier ones.
That's my experience as well. Gemini doesn't seem interested in doing searches outside of Deep Research mode, which is kind of funny given it should have the easiest access to a top search engine.
Re: Gemini 3.0 spotted in the wild through A/B testing
#149I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
This commonly expressed non-sequitur needs to die.
First of all, all of the big AI labs have crawled the internet. That's not a special advantage to Google.
Second, that's not even how modern LLMs are trained. That stopped with GPT-4. Now a lot more attention is paid to the quality of the training data. Intuitively, this makes sense. If you train the model on a lot of garbage examples, it will generate output of similar quality.
So, no, Google's crawling prowess has little to do with how good Gemini can be.
Re: Gemini 3.0 spotted in the wild through A/B testing
#150There are a lot more of these Gemini 3 examples out on twitter right now. After seeing them, I bought Google stock. What shocks me about its output is it actually feels like it's producing net new creative designs, not just regurgitated template output. Its extremely hard to design in code in a way that produces consistent, beautiful output, but it seems to be achieving it. That combined with Google being the only on…
But you do you if you have "fun money" to throw around!