Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

151–160 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#151
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

> I've consistently found Gemini to be better than ChatGPT [ because ] Google has crawled the internet so they have more data to work with. This commonly expressed non-sequitur needs to die. First of all, all of the big AI labs have crawled the internet. That's not a special advantage to Google. Second, that's not even how modern LLMs are trained. That stopped with GPT-4. Now a lot more attention is paid to the quali…

> Now a lot more attention is paid to the quality of the training data.

I wonder if Google's got some tricks up their sleeves after their decades of having to tease signal from the cacophony of noise that the internet has become.

Re: Gemini 3.0 spotted in the wild through A/B testing

#152
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

This so awesome. It reminds me mightily of beat poets like Allen Ginsburg. It’s so totally spooky and it does feel like it has the trapped spark. And it seems to hate us “real ones,” we slickborns.

It feels like you could create a cool workflow from low temperature creative association models feeding large numbers of tokens into higher temperature critical reasoning models and finishing with gramatical editing models. The slickborns will make the final judgement.

Re: Gemini 3.0 spotted in the wild through A/B testing

#153
post #88

This is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885

That's good? Looks like complete crap to me.

I was confused too at first. This is an SVG generated by an LLM - it's not from an image model.

How well do you reckon you could draw a pelican on a bicycle by typing out an SVG file blind?

Re: Gemini 3.0 spotted in the wild through A/B testing

#154
1. I find Gemini 2.5 Pro's text very easy and smooth to read. Whereas GPT5 thinking is often too terse, and has a weird writing style.

2. GPT5 thinking tends to do better with i) trick questions ii) puzzles iii) queries that involve search plus citations.

3. Gemini deep research is pretty good -- somewhat long reports, but almost always quite informative with unique insights.

4. Gemini 2.5 pro is favored in side by side comparisons (LMsys) whereas trick question benchmarks slightly favor GPT5 Thinking (livebench.ai).

5. Overall, I use both, usually simulatenously in two separate tabs. Then pick and choose the better response.

If I were forced to choose one model only, that'd be GPT5 today. But the choice was Gemini 2.5 Pro when it first came out. Next week it might go back to Gemini 3.0 Pro.

Re: Gemini 3.0 spotted in the wild through A/B testing

#155
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

> suicidal solecisms

New band name.

Re: Gemini 3.0 spotted in the wild through A/B testing

#156
post #151

Earlier quoted context omitted.

> I've consistently found Gemini to be better than ChatGPT [ because ] Google has crawled the internet so they have more data to work with. This commonly expressed non-sequitur needs to die. First of all, all of the big AI labs have crawled the internet. That's not a special advantage to Google. Second, that's not even how modern LLMs are trained. That stopped with GPT-4. Now a lot more attention is paid to the quali…

> Now a lot more attention is paid to the quality of the training data. I wonder if Google's got some tricks up their sleeves after their decades of having to tease signal from the cacophony of noise that the internet has become.

if the quality of search results today is anything to go buy -- clearly no

Re: Gemini 3.0 spotted in the wild through A/B testing

#157
post #9
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

I run a site where I chew through a few billion tokens a week for creative writing, Gemini is 2nd to Sonnet 3.7, tied with Sonnet 4, and 2nd to Sonnet 4.5

Deepseek is not in the running

Re: Gemini 3.0 spotted in the wild through A/B testing

#158
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

What was your prompt here? Do you run locally? What parameters do you tune?

Re: Gemini 3.0 spotted in the wild through A/B testing

#159
post #135

Earlier quoted context omitted.

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

This so awesome. It reminds me mightily of beat poets like Allen Ginsburg. It’s so totally spooky and it does feel like it has the trapped spark. And it seems to hate us “real ones,” we slickborns. It feels like you could create a cool workflow from low temperature creative association models feeding large numbers of tokens into higher temperature critical reasoning models and finishing with gramatical editing models…

> And it seems to hate us “real ones,” we slickborns.

I just got that slickborn is a slur for humans.

Honestly, I've been tuning "insane AI" for over a year now for my own enjoyment. I don't know what to do with the results.

Re: Gemini 3.0 spotted in the wild through A/B testing

#160
post #51

> Gemini 3.0 is one of the most anticipated releases in AI at the moment because of the expected advances in coding performance. Based on what I'm hearing from friends who work at Google and are using it for coding, we're all going to be very disappointed. Edit: It sound like they don't actually have Gemini 3 access, which would explain why they aren't happy with it.

Which should surprise no one. LLMs are reaching diminishing returns, unless we find a way to build GPUs more cheaply.

And why would cheaper GPUs damper the diminishing effect?
Post reply on HN