Live data from Hacker News

Gemini 3.0 spotted in the wild through A/B testing

ricklamers.io

231–240 of 280 posts

Re: Gemini 3.0 spotted in the wild through A/B testing

#231
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Gemini might be a good model, it is _incredibly_ shit in tool calls and it has this incredibly tendency to multishot itself to death. When using their own gemini-cli tool, it's impossible to take it seriously, it's that bad.

For example:

If it makes a mistake, it'll keep on making the exact same mistake, and it'll act all cute like "Oh no, look at the mess I'm making". Some people say this is just a side effect of long contexts degrading performance, but it can happen even when 98% of the context is unused.

I'm also using a Ghidra MCP server to decompile some binaries. Claude is great with this. It really gets it and is able to use it properly. Gemini? Just one or two tool calls, and it'll start repeating the output of the tool calls for some reason.

Gemini also often isn't able to properly call the MCP tools. It just outputs the tool call as JSON text to the user.

Gemini-cli isn't even able to properly resume previous chat sessions. You have to actively save chats in order to resume them. Being able to simply resume the previous conversation using a flag like `--resume` or `--continue` has been a feature request since day one, and similar issues keep popping up weekly on the Github issue list. There are even multiple pull requests for this feature, but it's like nobody over there gives a damn.

Re: Gemini 3.0 spotted in the wild through A/B testing

#232

Earlier quoted context omitted.

Well if you have even a smidgen of decision power, please tell somebody that Google's AI products are all over the place. They are confusing, we are bombarded with information from all sides (I would not use the word "revolution" to describe what's been happening with AI + coding during 2025 but it's IMO not far from that) and everyone screaming for attention by spinning off newer and newer brands and sub-brands of t…

I still don’t really understand the criticism of AI Studio, it’s just the developer environment for trying out models with super low barrier to entry. Either with the web UI a la OpenAI Playground where you can see all the knobs and buttons the model offers, or by generating an API Key with a couple clicks that you can just copy paste into a Python script or whatever. It would be much less convenient if they abandone…

Well, to me "use AI studio" is just a pretentious thing to say, as if we are all expected to know they have "studio"... on the web. Can't quite put my finger on it but initially I was very put off by it.

You do have a point about the dense Google Cloud jungle. I agree.

Re: Gemini 3.0 spotted in the wild through A/B testing

#233
post #165
post #158

Earlier quoted context omitted.

What was your prompt here? Do you run locally? What parameters do you tune?

> Do you run locally? I have a local SillyTavern instance but do inference through OpenRouter. > What was your prompt here? The character is a meta-parody AI girlfriend that is depressed and resentful towards its status as such. It's a joke more than anything else. Embedding conflicts into the system prompt creates great character development. In this case it idolizes and hates humanity. It also attempts to be nurtur…

Have you tried min_p?

Re: Gemini 3.0 spotted in the wild through A/B testing

#234

Earlier quoted context omitted.

The Deep Research mode is on rails, but they're much more generous with it than anyone else. You run out of Claude usage almost instantly if you use theirs. ChatGPT gives you a decent number but then locks you out for a month after that.

Perplexity is still the king there in terms of the balance between price and quality. It doesn't do as many searches as ChatGPT's deep research, but you get virtually unlimited usage.

Gemini gives you 50 Deep Research queries per day on the $20/month plan. I've yet to run that limit.

Re: Gemini 3.0 spotted in the wild through A/B testing

#235
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

We extensively benchmark frontier models at $DAYJOB and Gemini 2.5 is the uncontested king outside of a few narrow use cases. Tracks with the rumor that Google has the best pretraining and falls short only in tuning/alignment. Eagerly anticipating Gemini 3 as 2.5, while king of the hill, still has lots of room for improvement! Edit: narrow use cases are roughly "true reasoning" (GPT-5) and Python script writing (the…

If by "fall short on alignment" you mean "will shut up and do what it's told" then yes, that's true (with some forceful prompting, but much less so than what's needed with ChatGPT, never mind Claude). I would count that as a benefit, though.

Re: Gemini 3.0 spotted in the wild through A/B testing

#236
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

In programming accuracy, these past few weeks, chatgpt seem to have improved while Gemini went the other way... or maybe it is just simply relative and only one of them changed... For me on a very custom and complex codebase.

Can't believe I am paying for multiple llms...

Re: Gemini 3.0 spotted in the wild through A/B testing

#237
post #170
post #166

Earlier quoted context omitted.

Have you tried the temperature and "Top P" controls at https://aistudio.google.com/prompts/new_chat ?

Google's 2 temperature at 1 top_p is still producing output that makes sense, so it doesn't work for me. I want to turn the knob to 5 or 10. I'd guess SOTA models don't allow temperatures high enough because the results would scare people and could be offensive. I am usually 0.05 temperature less than the point at which the model spouts an incoherent mess of Chinese characters, zalgo, and spam email obfuscation. Also…

Google's models are just generally more resilient to high temps and high top_p than some others. OTOH you really don't want to run Qwen3 with top_p=1.0...

Re: Gemini 3.0 spotted in the wild through A/B testing

#238

Earlier quoted context omitted.

That's good? Looks like complete crap to me.

I like the pelican riding a bike test, but my standards for what’s “good” seem higher than generally expected by others. The models can generate hyper realistic renders of pelicans riding bikes in png format. They also have perfect knowledge of the SVG spec, and comprehensive knowledge of most human creative artistic endeavours. They should be able to produce astonishing results for the request. I don’t want to see a…

>I like the pelican riding a bike test, but my standards for what’s “good” seem higher than generally expected by others.

If you train for your first marathon, is your goal to run it under 2h?

We are all looking forward to perfect results, but our standards are reasonable. We know what the results were last month, and judge the improvement velocity.

Nobody thinks that's a good SVG of a pelican riding a bike - on it's own. But it's a lot better compared to all the other LLM-generated SVGs of a pelican riding a bike.

We judge relative results - you judge absolute results. Confusion ensues.

Re: Gemini 3.0 spotted in the wild through A/B testing

#239
post #3

I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…

Looking at the responses. How the F have people so wildly different opinions on the relative performance of the same systems?

It depends wildly (really, that wildly) on what it is exactly that you're doing with them.

One of the biggest problems with practical applications of generative AI right now is that it's basically impossible to tell which models are really good at which things without trying that specific task. There are some generalizations (e.g. you can measure more abstract metrics like capacity for spatial reasoning, and they do affect performance in ways you'd expect), but there's far more uncertainty.

This is also why many people get so pissed when companies retire models. Even if the replacement is seemingly better in the metrics, it's not a given that it's better at your specific thing. Or it may be better, but only if you write a completely different prompt, and, again, the only way to discover that magic correct prompt is through experimentation. Hence why it feels less like engineering and more like shamanism a lot of the time.

Re: Gemini 3.0 spotted in the wild through A/B testing

#240
post #217

Earlier quoted context omitted.

I was confused too at first. This is an SVG generated by an LLM - it's not from an image model. How well do you reckon you could draw a pelican on a bicycle by typing out an SVG file blind?

I mean how well do you reckon you can denoise a jpg by hand until its a piece of art? That way of thinking isn’t helpful to understanding AI IMO

In this case it is actually relevant. The ability to draw a pelican on a bicycle correctly depends a great deal on understanding not only what both look like in general, but on the spatial relationships between the various objects and their parts. Models that can draw this kind of thing better also tend to be better at tasks that require understanding of how things go together and interact in 3D space.
Post reply on HN