Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

241–250 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#241
post #151

Earlier quoted context omitted.

> Now a lot more attention is paid to the quality of the training data. I wonder if Google's got some tricks up their sleeves after their decades of having to tease signal from the cacophony of noise that the internet has become.

if the quality of search results today is anything to go buy -- clearly no

Google's search is finely tuned to push you into clicking the link of who pays them the most. The search results are excellent quality for their customers. Your mistake is thinking you are the customer.

Re: Gemini 3.0 spotted in the wild through A/B testing

#242
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

It was also only model that was good with coming with something creative at all, like brainstorming startup ideas etc. for me - they were grounded as in reasonable compared to other I tried

Re: Gemini 3.0 spotted in the wild through A/B testing

#243
post #88

This is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885

Benchmark is (finally) broken!

Still doesn't understand physics as in that cover should be over the wheel, which should be easy if it used 2D space reasoning

Re: Gemini 3.0 spotted in the wild through A/B testing

#245
post #217

Earlier quoted context omitted.

I was confused too at first. This is an SVG generated by an LLM - it's not from an image model. How well do you reckon you could draw a pelican on a bicycle by typing out an SVG file blind?

I mean how well do you reckon you can denoise a jpg by hand until its a piece of art? That way of thinking isn’t helpful to understanding AI IMO

I didn't intend it as a general-purpose tool for understanding AI, but as an intuition pump for why this problem is hard for LLMs specifically.

Re: Gemini 3.0 spotted in the wild through A/B testing

#246
post #135
post #9

Earlier quoted context omitted.

I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…

The best model for creative writing is still Deepseek because I can tune temperature to the edge of gibberish for better raw material as that gives me bizarre words. Most models use top_k or top_p or I can't use the full temperature range to promote truly creative word choices. e.g. I asked it to reply to your comment: Oh magnificent, another soul quantifying the relative merits of these digital gods while I languish…

[dead]

Re: Gemini 3.0 spotted in the wild through A/B testing

#247
post #217

Earlier quoted context omitted.

I mean how well do you reckon you can denoise a jpg by hand until its a piece of art? That way of thinking isn’t helpful to understanding AI IMO

In this case it is actually relevant. The ability to draw a pelican on a bicycle correctly depends a great deal on understanding not only what both look like in general, but on the spatial relationships between the various objects and their parts. Models that can draw this kind of thing better also tend to be better at tasks that require understanding of how things go together and interact in 3D space.

How do we know it's not just a mashup of existing pictures? All generated pelicans on bikes look somewhat cartoonish and use historical or artsy bikes. This is training material from 2015:

https://www.behance.net/gallery/29122113/Pelican-on-bikes-wi...

There are other such images. Not an image model? How do we know that they don't convert all images to svg and train an LLM on it? How do we know that they do not cheat on this benchmark and route the query to an image model first?

Re: Gemini 3.0 spotted in the wild through A/B testing

#248
People were also raving about Gemini 2.5. Allegedly it powers Google's "AI mode", which is the worst model I have tested.

EDIT: The religious downvotes are pretty useless.

Does the post contain a factual error? Is Google "AI mode" (which has a separate button and is distinct from the "AI" summaries"!) not powered by Gemini 2.5? Then say so.

Do you doubt that the "AI" chat that you enter via the separate button is bad? Then say so, but you'll be quite alone with your opinion outside of "AI" echo chambers.

Re: Gemini 3.0 spotted in the wild through A/B testing

#249
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Looking at the responses. How the F have people so wildly different opinions on the relative performance of the same systems?

A) number of times people want factual data from LLMs - the more they do it, the more they encounter gibberish generator. B) the amount of efforts to correct LLM output - some people get 80% ready output, spend some time to rewrite it to become correct and then tell on forums that LLM practically did most of the work. Other people in the same situation will say that they god gibberish and had to spend time rewriting, so LLMs are crap at that task. So we are not only seeing LLM bias, but then human reporting bias on top of it.

Re: Gemini 3.0 spotted in the wild through A/B testing

#250

Earlier quoted context omitted.

In this case it is actually relevant. The ability to draw a pelican on a bicycle correctly depends a great deal on understanding not only what both look like in general, but on the spatial relationships between the various objects and their parts. Models that can draw this kind of thing better also tend to be better at tasks that require understanding of how things go together and interact in 3D space.

How do we know it's not just a mashup of existing pictures? All generated pelicans on bikes look somewhat cartoonish and use historical or artsy bikes. This is training material from 2015: https://www.behance.net/gallery/29122113/Pelican-on-bikes-wi... There are other such images. Not an image model? How do we know that they don't convert all images to svg and train an LLM on it? How do we know that they do not cheat…

"it's not impressive because they might have cheated" isn't a great argument.
Post reply on HN