Earlier quoted context omitted.
I started trying out 3.7 Flash this week and it is competitive with opus/fable and also FAST. It is getting work done that anthropic models were struggling with and the speed with which it does is quite a bit noticeably faster. Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.
I use Gemini because I feel like Google will win the AI race, and it’s Good Enough
Gemini 3.8 Flash and 3.8 Flash Cyber
551–560 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#552Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#553Earlier quoted context omitted.
Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing ! I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token. Here is the result https://coolthing-ling-3-tiny.tiiny.site (sorry for the weird hosting, first I found that worked) (it cost…
> Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing ! Datasets contains lots of people sharing particles simulations in various ways, with a bunch of people replying "that's so cool" and similar, so 10 years later someone asks an LLM for "cool thing" and "particle simulations" rank pretty far up when it thinks about what others have call…
Also the reason LLMs are positive, enchanting, pleasant, glorifying, demagogues.
Not because it's skewed tone in the data. They are acute politicians.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#554Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
The result is the same as the previous Gemini 3.6/3.7 Flash days: Claude could always note much more problems in Codex's plan and implementation than Gemini could - the ratio is like 10:1.
I occasionally switch the roles between Codex and Claude, and result is the same, Codex could always catch much more problems in Claude's plan and implementation, than Gemini could.
So I am guessing in a relatedly complex codebase, Gemini is much less effective in acting as a guardrail (or a senior engineer/team lead) than the other SOTA models.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#555Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#556Earlier quoted context omitted.
Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…
Adding another agent to check the first one feels like putting a band-aid on a band-aid. If there is an issue with the third one, we adding a fourth one as well
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#557Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#558Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#559Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability…
Gemini is known for good at creative writing in the Chinese writing community. It's a bit ironic though. Google has probably the most and best code base among all tech companies but its Gemini is bad at coding. Google has no access to Chinese market but its model is incredibly good at writing in Chinese.
probably because google knows people would want to extract google's proprietary code from gemini if they use their own code base to train it! I bet they purposefully gimped it to prevent that from happening.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#560Earlier quoted context omitted.
That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.
Easy, have another agent check it. Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues. The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.
[Claude proceeds to waste your time telling you about bugs it caused then fixed and other non-issues...]
Really wish they'd get rid of this. It must be in the system prompt as it always 'flags' 2 things