Beats 3.1 Pro for price per token, but artificial analysis is showing it's dumber per token and costs more overall
Gemini 3.5 Flash
71–80 of 692 posts
Re: Gemini 3.5 Flash
#72Re: Gemini 3.5 Flash
#73Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
People complain about them incessantly, but I can almost never get people to actually post receipts. Every provider allows sharing chats, and anyone can share a prompt that reliably produces hallucinations. More often than not, people are using images in responses that go awry. Which is fair, the models are sold as multi-modal, but image analyses is still at gpt-4.0 text-analyses levels. Also knowledge cutoff issues,…
And when I say all the time, I mean it, and this is for Opus 4.7 Adaptive.
I often have to say, please do searches and cite sources, as if it doesn't it will confidently give me wrong or outdated information.
If you're often asking questions about a topic that's not in your specialist knowledge you won't notice them.
Re: Gemini 3.5 Flash
#74Earlier quoted context omitted.
I think they mean the boat is moving. In the flash ones the paddles are animated but the boat is stationary for me.
The boat moves in all three for me
Re: Gemini 3.5 Flash
#75It’s not possible to uptrain on preview releases and it did not get that much love for a while.
Re: Gemini 3.5 Flash
#76Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
if last year's models were the ones people got familiar with in late 2022, hallucinations would be an underrepresented rumor, there would be no articles about it because its so rare. overconfident lawyers wouldn't have messed up dockets in court with fake case law, in other domains that move faster, sources would be only partially outdated with agentic search and mcp servers filling in the gaps AI psychosis would be…
Re: Gemini 3.5 Flash
#77Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
Re: Gemini 3.5 Flash
#783.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite. $1551 for 3.5 Flash [0] vs $892 for 3.1 Pro [1]. That's 74% more cost while ranking lower. It's 2.5x as fast but I don't think the bang for the buck is there anymore like it was with 3.0 Flash. I'm a bit bummed out to be honest. I did not expect such a huge (3x) price increase from 3.0 Flash and I bet many people will not just blindly…
Does that mean this model is production ready?
Re: Gemini 3.5 Flash
#79Plus the vibe of the gemini models are so weird particularly when it comes to tool calling
At this point I kinda need them to shock me to make the switch
Re: Gemini 3.5 Flash
#80Also concerned about Gemini models being benchmaxxed generally