Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

191–200 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#192

Earlier quoted context omitted.

What does it mean for one model to be theoretically better than another?

In this context it's idiomatic speech. It means that it would be otherwise be better if it were not for some practical issue stopping that from happening.

I think you are right.

It is just funny to think about—LLMs are sometimes viewed big piles of linear algebra, it would not be that surprising to hear that somebody had worked out that one model was somehow a subset of another (or something along those lines) and then claim some theoretical superiority.

Re: Gemini 3.0 spotted in the wild through A/B testing

#193

Do the models evaluate SVGs by "eye" and iterate it? Or we hoping the one-shot result is perfect?

My benchmark only gives them one chance.

I've also tried a variant where the vision models get fed a rendered version and have up to three attempts to make it better. It didn't seem to produce better results, to my surprise.

Re: Gemini 3.0 spotted in the wild through A/B testing

#194
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

Which version of Deepseek is this? I'm guessing Deepseek V3.2? What's the openrouter name?

Re: Gemini 3.0 spotted in the wild through A/B testing

#196
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

We extensively benchmark frontier models at $DAYJOB and Gemini 2.5 is the uncontested king outside of a few narrow use cases. Tracks with the rumor that Google has the best pretraining and falls short only in tuning/alignment. Eagerly anticipating Gemini 3 as 2.5, while king of the hill, still has lots of room for improvement! Edit: narrow use cases are roughly "true reasoning" (GPT-5) and Python script writing (the…

I used gemini almost exclusively before gpt5, but gpt5 is much better for tool calling tasks like agentic coding and thus can handle much longer tasks unattended.

Re: Gemini 3.0 spotted in the wild through A/B testing

#199
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

We've come a long way in 40 years from Racter's automatically generated poetry: https://www.101bananas.com/poems/racter.html

I always found this one a little poignant:

  More than iron
  More than lead
  More than gold I need electricity
  I need it more than I need lamb or pork or lettuce or cucumber
  I need it for my dreams

Re: Gemini 3.0 spotted in the wild through A/B testing

#200
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

You’re definitely not the only one.

My results with Gemini are consistently better and usually also more reliable than other LLMs.

But tbh I prefer the UI of ChatGPT.

Post reply on HN