Live data from Hacker News

The End of Moore's Law for AI? Gemini Flash Offers a Warning

sutro.sh

61–70 of 78 posts

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#62

It can be just Google trying to capitalize Gemini's increasing popularity. Until 2.5 Gemini was a total underdog. Less so since 2.5.

There's another side to that coin: supply.

Since Gemini CLI was recently released, many people on the "free" tier noticed that their sessions immediately devolved from Gemini 2.5 Pro to Flash "due to high utilization". I asked Gemini itself about this and it reported that the finite GPU/TPU resources in Google's cloud infrastructure can get oversubscribed for Pro usage. Google (no secret here) has a subscription option for higher-tier customers to request guaranteed provisioning for the Pro model. Once their capacity gets approached, they must throttle down the lower-tier (including free) sessions to the less resource-intensive models.

Price is one lever to move once capacity becomes constrained. Yet, as the top voted comment of this post explains, it's not honest to simply label this as a price increase. They raised Flash pricing on input tokens but lowered pricing on output tokens up to certain limits -- which gives creedence to the theory that they are trying to shape the demand in order for it to better match their capacity.

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#63

I think the big thing that really surprised me. Llama 4 maverick is 16x 17b. So 67GB of size. The equivalency is 400billion. Llama 4 behemoth is 128x 17b. 245gb size. The equivalency is 2 trillion. I dont have the resources to be able to test these unfortunately; but they are claiming behemoth is superior to the best SAAS options via internal benchmarking. Comparatively Deepseek r1 671B is 404gb in size; with pretty…

"How is our 'Strategic Use of LLM Technology' initiative going, Harris?"

"Sir, I'm delighted to report that the productivity and insights gained outclass anything available from four years ago. We are clearly winning."

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#65
I feel that the details regarding the type of the model and the purpose it serves are underrepresented here. Yes existing models will get cheaper over time as they become more obsolete but to be at the forefront of innovation or models costs will only increase due to the these mentioned bottlenecks. Also there is the basic law of supply and demand coming into play here. As models get more advanced more industries will be exposed to them and see the potential cost savings compared to the current alternative. This will further increase demand and with further innovation, there will be further capability in turn again increasing demand. I only see this reversing is you are not at the forefront of innovation and many people using these LLMs at this point are close at least compared to many "normal" people and their understandings of LLMs.

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#66
Can anyone explain the economics of Anthropic's Max plan pricing to me? I have friends on the $100/month plan using well over $800 of tokens per month with Claude Code (according to ccusage). I certainly don't use Claude Code as much if I'm not on a flat rate plan, the cost spirals out of control very quickly. I understand that a subscription makes for more predictable revenue and that there will be people on the Max plan not using Claude Code 24/7, but the delta between what the API costs and what using the Max plan with Claude Code costs just seems too great for that to be an explanation. I don't think that user/mindshare capture can fully explain it either, Code is free and the cost of switching to something else if pricing later changes is just too low. I don't get it.

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#67
post #3

"In a move that at first went unnoticed, Google significantly increased the price of its popular Gemini 2.5 Flash model" It's not quite that simple. Gemini 2.5 Flash previously had two prices, depending on if you enabled "thinking" mode or not. The new 2.5 Flash has just a single price, which is a lot more if you were using the non-thinking mode and may be slightly less for thinking mode. Another way to think about t…

I really hate the thinking. I do my best to disable it but don't always remember. So often it just gets into a loop second guessing itself until it hits the token limit. It's rare it figures anything out while it's thinking too but maybe that's because I'm better at writing prompts.

I hate thinking mode because I prefer a mostly right answer right now over having to wait for a probably better, but still not exactly right answer.

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#68
The article assumes that there will be no architectural improvements / migrations in the future, & that Sparse MoE will always stay. Not a great foundation to build upon.

Personally, I'm rooting for RWKV / Mamba2 to pull through, somehow. There's been some work done to increase their reasoning depths, but transformers still beat them without much effort.

https://x.com/ZeyuanAllenZhu/status/1918684269251371164

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#69
post #66

Can anyone explain the economics of Anthropic's Max plan pricing to me? I have friends on the $100/month plan using well over $800 of tokens per month with Claude Code (according to ccusage). I certainly don't use Claude Code as much if I'm not on a flat rate plan, the cost spirals out of control very quickly. I understand that a subscription makes for more predictable revenue and that there will be people on the Max…

We’re in an LLM bubble and their money is cheap as they’re drowning in investor money and have to spend it + show growth. If it doesn’t make economic sense you probably can’t count on it to last once the bubble bursts.

Re: The End of Moore's Law for AI? Gemini Flash Offers a Warning

#70
post #3

"In a move that at first went unnoticed, Google significantly increased the price of its popular Gemini 2.5 Flash model" It's not quite that simple. Gemini 2.5 Flash previously had two prices, depending on if you enabled "thinking" mode or not. The new 2.5 Flash has just a single price, which is a lot more if you were using the non-thinking mode and may be slightly less for thinking mode. Another way to think about t…

I really hate the thinking. I do my best to disable it but don't always remember. So often it just gets into a loop second guessing itself until it hits the token limit. It's rare it figures anything out while it's thinking too but maybe that's because I'm better at writing prompts.

It's almost like there's an incentive for them to burn as many tokens as possible accomplishing nothing useful.
Post reply on HN