Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
Gemini 3.8 Flash and 3.8 Flash Cyber
51–60 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#52Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC? I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker! At this point it is a meme of course, but where is 3.5 Pro :)
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#53But where is Gemini 3.5 pro?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#54Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.
Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#55Looks like Google's given up on frontier models for external consumption?
Latest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#56Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#57Earlier quoted context omitted.
IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High. But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny". Flash is my go-to for prototyping, and basically anything that isn't writing production code.
The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#58Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#59And yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#60Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.
So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.