Gemini 3.0 spotted in the wild through A/B testing
191–200 of 280 posts
Re: Gemini 3.0 spotted in the wild through A/B testing
#192Earlier quoted context omitted.
What does it mean for one model to be theoretically better than another?
In this context it's idiomatic speech. It means that it would be otherwise be better if it were not for some practical issue stopping that from happening.
It is just funny to think about—LLMs are sometimes viewed big piles of linear algebra, it would not be that surprising to hear that somebody had worked out that one model was somehow a subset of another (or something along those lines) and then claim some theoretical superiority.
Re: Gemini 3.0 spotted in the wild through A/B testing
#193Do the models evaluate SVGs by "eye" and iterate it? Or we hoping the one-shot result is perfect?
I've also tried a variant where the vision models get fed a rendered version and have up to three attempts to make it better. It didn't seem to produce better results, to my surprise.
Re: Gemini 3.0 spotted in the wild through A/B testing
#194Earlier quoted context omitted.
I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…
The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…
Re: Gemini 3.0 spotted in the wild through A/B testing
#195Re: Gemini 3.0 spotted in the wild through A/B testing
#196I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
We extensively benchmark frontier models at $DAYJOB and Gemini 2.5 is the uncontested king outside of a few narrow use cases. Tracks with the rumor that Google has the best pretraining and falls short only in tuning/alignment. Eagerly anticipating Gemini 3 as 2.5, while king of the hill, still has lots of room for improvement! Edit: narrow use cases are roughly "true reasoning" (GPT-5) and Python script writing (the…
Re: Gemini 3.0 spotted in the wild through A/B testing
#197My friends at Google hate AI coding with passion. I have some theories as to why. But anyone here venture a guess?
Re: Gemini 3.0 spotted in the wild through A/B testing
#198My friends at Google hate AI coding with passion. I have some theories as to why. But anyone here venture a guess?
Re: Gemini 3.0 spotted in the wild through A/B testing
#199Earlier quoted context omitted.
I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…
The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…
I always found this one a little poignant:
More than iron
More than lead
More than gold I need electricity
I need it more than I need lamb or pork or lettuce or cucumber
I need it for my dreamsRe: Gemini 3.0 spotted in the wild through A/B testing
#200I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
My results with Gemini are consistently better and usually also more reliable than other LLMs.
But tbh I prefer the UI of ChatGPT.