Earlier quoted context omitted.
Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else. Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?
I use Gemini every day and I've noticed any subjective improvement. In many cases it feels worse because it does fewer Google searches than before. As a result I find it hard to get excited about it. I do think Gemini is underrated on HN though!
Gemini 3.8 Flash and 3.8 Flash Cyber
621–630 of 700 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#622Earlier quoted context omitted.
the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4. 3.0 flash -> 3.8 flash is all post training which is pretty impressive.
Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.
AI models are almost completely interchangeable, so the best/cheapest/fastest whatever will always have a market.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#623I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…
Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#624I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…
> Real world knowledge For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.
That being said, Anthropic is also so insanely expensive for everything I ended up switching that particular part to Exo instead...
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#625I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…
Never had a bad suggestion.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#626Earlier quoted context omitted.
I setup Luna as main Claude Code driver (so zero anthropic api use) and it nailed crisply a handful of python tasks, gonna continue this way.
Why not use codex or an open source harness?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#627The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting. Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html": https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
it's such a weird split how most AI companies are trying to be the best, but Google really has a different mission statement. they already have users. lots of users. they need to be working on building models they can deploy and use with the most number of people, as they already have the users. i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-l…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#628Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents (I think thinking level low is a regression on 3.8 compared to 3.7.)
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#629Earlier quoted context omitted.
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis. Lu…
I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.
I also included Sol56 xhigh, which ranks above even Gemini38 high.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#630Earlier quoted context omitted.
That's the thing. I am completely lost because there are so many redundant paths to get the same thing and I'm trying to figure out which one is the best deal
Well are you looking for a subscription or pay-as-you-go API usage? Subscription? -> Google One plan (http://one.google.com/) API? -> AI Studio (https://aistudio.google.com/) It's not really any different than the choice you'd make with OpenAI/Anthropic depending on how you plan to use it. Except as a hyperscalar, it's also offered first party from Google Cloud (like Claude via Amazon Bedrock or GPT via Microsoft Azu…
I have spent two months experimenting with a wide range of US and Chinese models, and I had a lot of fun doing that, but I am in the process of switching to just using local models, using Gemini on an API if I need it, and once or twice a month when I really need help on something difficult, I use something top-tier like Kimi K3.