This is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885
That's good? Looks like complete crap to me.
Gemini 3.0 spotted in the wild through A/B testing
111–120 of 280 posts
Re: Gemini 3.0 spotted in the wild through A/B testing
#112Earlier quoted context omitted.
Gemini is theoretically better, but I find it's very unsteerable. Combine that with the fact it struggles with tool use and character-level issues - and it can be challenging to use despite being "smarter".
What does it mean for one model to be theoretically better than another?
Re: Gemini 3.0 spotted in the wild through A/B testing
#113This is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885
That's good? Looks like complete crap to me.
The models can generate hyper realistic renders of pelicans riding bikes in png format. They also have perfect knowledge of the SVG spec, and comprehensive knowledge of most human creative artistic endeavours. They should be able to produce astonishing results for the request.
I don’t want to see a chunky icon-styled vector graphic. I want to see one of these models meticulously paint what is unambiguously a pelican riding what is unambiguously a bicycle, to a quality on-par with Michelangelo, using the SVG standard as a medium. And I don’t just want it to define individual pixels. I want brush strokes building up a layered and textured birds wing.
Re: Gemini 3.0 spotted in the wild through A/B testing
#114Earlier quoted context omitted.
using an LLM for "creative writing" is like getting on a motorcycle and then claiming you went for a ride on a bicycle no, wait, that analogy isn't even right. it's like going to watch a marathon and then claiming you ran in it.
Just imagine you’re trying to build a custom D&D campaign for your friends. You might have a fun idea don’t have the time or skills to write yourself that you can have an LLM help out with. Or at least make a first draft you can run with. What do your friends care if you wrote it yourself or used an LLM? The quality bar is going to be fairly low either way, and if it provides some variation from the typical story boo…
If I found out a player had come to the table with an LLM generated character, I would feel a pretty big betrayal of trust. It doesn't matter to me how "good" or "polished" their ideas are, what matters is that they are their own.
Similarly, I would be betraying my players by using an LLM to generate content for our shared game. I'm not just an officiant of rules, I'm participating in shared storytelling.
I'm sure there are people who play DnD for reasons other than storytelling, and I'm totally fine with that. But for storytelling in particular, I think LLM content is a terrible idea.
Re: Gemini 3.0 spotted in the wild through A/B testing
#115I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
Re: Gemini 3.0 spotted in the wild through A/B testing
#116I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
I agree with you, I consistently find Gemini 2.5 Pro better than Claude and GPT-5 for the following cases: * Creative writing: Gemini is the unmatched winner here by a huge margin. I would personally go so far as to say Gemini 2.5 Pro is the only borderline kinda-sorta usable model for creative writing if you squint your eyes. I use it to criticize my creative writing (poetry, short stories) and no other model unders…
Re: Gemini 3.0 spotted in the wild through A/B testing
#117My strange observation is that Gemini 2.5 Pro is maybe the best model overall for many use cases, but starting from the first chat. In other words, if it has all the context it needs and produces one output, it's excellent. The longer a chat goes, it gets worse very quickly. Which is strange because it has a much longer context window than other models. I have found a good way to use it is to drop the entire huge con…
> The longer a chat goes, it gets worse very quickly. This has been the same for every single LLM I've used, ever, they're all terrible at that. So terrible that I've stopped going beyond two messages in total. If it doesn't get it right at the first try, its more and more unlikely to get it right for every message you add. Better to always start fresh, iterate on the initial prompt instead.
Re: Gemini 3.0 spotted in the wild through A/B testing
#118Earlier quoted context omitted.
I gave up on Gemini because I couldn't stop the glazing. I don't need to be told what can incredible insight I have made and why my question gets to the heart of the matter every time I ask something.
With AI studio there's a system prompt where you can tell it to stop the sycophancy. But yeah it does do that otherwise. At one point it told me I'm a genius.
Re: Gemini 3.0 spotted in the wild through A/B testing
#119I might be in the minority here but I've consistently found Gemini to be better than ChatGPT, Claude and Deepseek (I get access to all of the pro models through work) Maybe it's just the kind of work I'm doing, a lot of web development with html/scss, and Google has crawled the internet so they have more data to work with. I reckon different models are better at different kinds of work, but Gemini is pretty excellent…
I've since switched to Claude Code and I no longer have to spend nearly as much time managing context and scope.
Re: Gemini 3.0 spotted in the wild through A/B testing
#120This is a very good pelican. I'm really looking forward to trying out Gemini 3 myself. https://x.com/cannn064/status/1978779247930953885
That's good? Looks like complete crap to me.