Earlier quoted context omitted.
I think it's good at basic coding and is cheaper and faster. Trying it vs Fable 5.1 I don't see it being even close capability-wise. What I mean by basic coding is "write a script to parse this data" and answering some questions based on the code. It does that well and answers fast. However, when it comes to planning and thinking through a more complicated design problem it's just not there yet.
It's not even close. Some benchmarks show 3.8 Flash as being close to Astra (DeepSWE v1.1), but they are mostly "bench-maxed". Newer benchmarks like Terminal-Bench 4.0 show Astra at 57.9% and 3.8 Flash at 19.1%. For anyone who has tried to code with Gemini models, there is no contest here. Astra wins handily on every other benchmark, too. Research, science, 3D, visual, Humanity's Last Exam, etc.
Gemini 3.8 Live and 3.8 Live Extended Thinking
331–334 of 334 posts
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#332I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.
I love it as a variation from the others. 3.8 flash is the best back-and-forth model for iterating imo, but would not use for long horizon
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#333Earlier quoted context omitted.
It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.
I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#334Earlier quoted context omitted.
I love it as a variation from the others. 3.8 flash is the best back-and-forth model for iterating imo, but would not use for long horizon
As I mentioned - I find it awful for iterating. My recent example: I was researching shoes for toddler, wide, boa or similar mechanism instead of laces/velcro. I did same prompt in ChatGPT and flash. After several back and forth responses/clarifications flash completely lost track of what I’m looking for and started suggesting nonsense.