Gemini 3.8 Flash and 3.8 Flash Cyber
641–650 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#642Earlier quoted context omitted.
Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.
What harness do you use for Gemini? Antigravity?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#643Earlier quoted context omitted.
As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash). With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3. https://imgur.com/a/BMOJBED
They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#644Is there any comparison of usage limits for Antigravity plans vs. Codex? I just ran two light tasks on my codebase and got 100% of the weekly limits of a Pro plan blown away. Is Ultra plan any different? Because on Codex it wouldn't affect my Max plan at all, I think it would have been below 1% othese usage.
Anthropic and OpenAI are in a different game of spending investor money to buy market share.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#645Earlier quoted context omitted.
Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.
Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#646I can’t wait until waiting hours and spending a big chunk of your usage per task seems antiquated, and real-time iteration on massive code changes is the norm. This might just be the year of efficiency, that truly allows AI to be used to the heart’s content.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#647The biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?
thats not just the context growing issue, hallucinations is a thing
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#648Earlier quoted context omitted.
sidenote, but wow sonnet 5 is shockingly bad on this benchmark.
sonnet 5 is bad by almost any metric. anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.
I find GLM5.3 so much better than Sonnet it is not even funny.
Sonnet behaves like a cheap model while being very expensive.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#649Earlier quoted context omitted.
Because intent supposes will which supposes consciousness, and these aren't.
I'll agree if you can define consciousness in a way that: 1) Excludes what LLM's do. 2) Doesn't exclude what many humans do (including the neuro divergent). 3) Doesn't just boil do to simply rephrasing your pre-existing belief/prejudice that humans are conscious and nothing else can be as if it were a fact and not an opinion. I suspect that you can't.