Live data from Hacker News

Gemini 3.5 Flash

blog.google

281–290 of 692 posts

Re: Gemini 3.5 Flash

#281
post #131

The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...

``

wtf

``

WTF??

Re: Gemini 3.5 Flash

#282

Knowledge cutoff: January 2025 Latest update: May 2026 I have a very bad feeling about this lag.

At least in some cases, there seems to be a move toward training on more synthetic data and strictly curated data, especially for smaller models where knowledge can't be extremely broad, because there just isn't enough room to store the world in tens or hundreds of gigabytes of model weights. So, to achieve higher quality reasoning, the training has to be focused and the data has to be very high quality and high dens…

> it maybe doesn't even matter that the models are using older data.

This actually really does matter. Otherwise, the model simply won't know about your product and will always suggest only a few market leaders.

Searching for information on the Internet became a jungle a decade ago, and to be visible you have to pay Google for sunlight. Now, we risk falling into real darkness — until some paid model eventually emerges. This might be the reason Google is fine with training data from 2024. If the top spot is reserved for whoever pays anyway, why bother?

Re: Gemini 3.5 Flash

#283

Earlier quoted context omitted.

To be honest, China not having access to the latest hardware is exactly what has driven LLM technology forward the last 2 years.

Why?

Because it forced them to focus on efficiency, instead of throwing more compute at the problem.

Just like in software, some of the most beautiful solutions come from constraints. Think, the optimisations that game developers implemented because of the frame budget.

Re: Gemini 3.5 Flash

#284
post #223

Am I really so old that when someone says "Flash" my immediate response is... "consider HTML5 instead" ??

Lol. Young uns! Flash, ah, ah, saviour of the universe. Flash, ah, ah, he'll save every one of us! Every time I have heard the word flash for goodness knows how many years.

If Google can reuse the "Flash" brand, I'm re-branding myself as "Meadhbh the Merciless."

Re: Gemini 3.5 Flash

#286

Earlier quoted context omitted.

These companies are unprofitable (as all companies at this stage and ambition should be) but I increasingly don't see any justification for the idea that it is fundamentally unprofitable. Inference alone is certainly profitable. I'm running models at home that are comparable to performance of paid models a year or so ago for free. Even for much larger models the cost around inference serving are clearly manageable. T…

If it's profitable, why haven't they reported any profits? People like Ed Zitron have done the math and it just doesn't add up. I mean he just published this piece today: https://www.wheresyoured.at/ai-is-too-expensive/

His entire brand is that the AI bubble will burst. By his account it was supposed to have several times by now. Like the doomers, it's not if it's when and they have to keep pushing back their predictions. Funny how both camps can be so confident. Alas, that's how they get eyes, ears and dollars.

That's not to say they will be or are wrong, it's just that they aren't exactly unbiased, or humble, sources.

Re: Gemini 3.5 Flash

#287

Earlier quoted context omitted.

We need another "Deepseek moment" or else it will become impossible for the regular dude to use AI. It will become something that only big companies can afford.

We're having DeepSeek moments every couple of weeks. Qwen 3.6 hit hard in the self-hosting space. It's incredibly capable for its size, really shaking up what's possible in 64GB or even 32GB of VRAM. The Prism Bonsai ternary model crams a tremendous amount of capability into 1.75GB. And, DeepSeek V4 is crazy good for the price. They're charging flash model prices for their top-tier Pro model, which is competitive wit…

> It's incredibly capable for its size, really shaking up what's possible in 64GB or even 32GB of VRAM.

You can lower that to at least 24GB. I've been running Qwen 3.5 and 3.6 with codex on a 7900 XTX and the long horizon tasks it can handle successfully has been blowing my mind. I would seriously choose running my current local setup over (the SOTA models + ecosystem) of a year ago just based on how productive I can be.

Re: Gemini 3.5 Flash

#288

Am I really so old that when someone says "Flash" my immediate response is... "consider HTML5 instead" ??

The Flash designer was really nice. One thing the web kind of set back was all the RAD tools from the 90s and 2000s.

And there were some amazing RAD and prototyping tools in the 90s (mostly for DOS, but also for Windoze desktop apps.) You're right, we sort of gave up on the idea when everyone wanted to be seen as a "real" software engineer who knew how to sling Java on the back end.

Re: Gemini 3.5 Flash

#289
post #265

I have google ai pro plan and tried antigravity with 3.5 flash but it used up all my quota in two prompts. If that is not a bug then it is seriously unusable.

Yesterday, or the day before, Google lowered the AI Pro quota from 33x standard usage to 4x.

From the talk on the Gemini subreddit it's severely lower than before. I'm likely canceling my AI Pro.

The update also broke the app for me. Editing a message crashes the app every time. I'm on a Pixel lol

Re: Gemini 3.5 Flash

#290

Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase for in the immediate next same-sized model (and lol @ 3 only ever getting a preview). 3.5 flash costs similar to Gemini 2.5 pro which was $1.25/$10

We need another "Deepseek moment" or else it will become impossible for the regular dude to use AI. It will become something that only big companies can afford.

Maybe we can figure out better ways to use the models that can run on cheap hardware.
Post reply on HN