Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

641–650 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#641

Earlier quoted context omitted.

Hi, thanks for Claude Code. I was wondering though if you'd considering adding a mode to make text green and characters come down from the top of the screen individually, like in The Matrix?

No, I want a little monkey doing tricks. /s

This guy? https://en.wikipedia.org/wiki/BonziBuddy

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#642

I'm noticing a fair number of degradation of Claude infrastructure recently and makes me wonder why they can't use Claude to identify or fix these issues in advance? It seems a counter intuitive to Anthropic's message that Claude uncovered bugs in open source project*. [*] https://www.anthropic.com/news/mozilla-firefox-security

timing wise this seems to match the Claude Mythos story.

So maybe they're trying to free-up some GPU capacity to run audit of projects in need? I'm assuming Mythos is not cheap to run.

The cache TTL story is also probably link to the RAM price going up like mad so they're trying to save on future expenditure here maybe?

I do understand why people are pissed though

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#643
post #589

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

So Anthropic is trying to save money on infrastructure, we all get it. However, it's not ok to degrade the performance your users have paid for. Last week the issue was that you reduced the default "effort" level, now the prompt cache is shortened. Several users experience far more restrictive usage limits lately. There is only so much you can do through "UX improvements" or some smart routing on the backend. Your fl…

For context, my company gives each developer a decent monthly allowance for Claude and if push comes to shove, we are allowed to fallback to using AWS Bedrock hosted Anthropic models.

When you pay for a Claude subscription, what exactly were you promised?

> they will start voting with their money.

And go where? Sooner or later the party is going to be over and Claude and its competitors are going to have to start charging enough to actually be profitable when the VC money dries up.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#644
post #496
post #487

Earlier quoted context omitted.

Where are you getting 60GB from? It shouldn’t be that large. But yes, would love to save context/cache such that it can be played back/referred to if needed. /compact is a little black box that I just have to trust that is keeping the important bits.

The KV cache consists of activation vectors for every attention head at every layer of the model for every token, so it gets quite large. ChatGPT also estimates 60-100GB for full token context of an Opus-sized model: https://chatgpt.com/share/69dc5030-268c-83e8-92c2-6cef962dc5...

That is actually nuts.... I'm trying to understand the true costs of AI, wonder how I plug this in!

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#645
post #589

Earlier quoted context omitted.

So Anthropic is trying to save money on infrastructure, we all get it. However, it's not ok to degrade the performance your users have paid for. Last week the issue was that you reduced the default "effort" level, now the prompt cache is shortened. Several users experience far more restrictive usage limits lately. There is only so much you can do through "UX improvements" or some smart routing on the backend. Your fl…

For context, my company gives each developer a decent monthly allowance for Claude and if push comes to shove, we are allowed to fallback to using AWS Bedrock hosted Anthropic models. When you pay for a Claude subscription, what exactly were you promised? > they will start voting with their money. And go where? Sooner or later the party is going to be over and Claude and its competitors are going to have to start cha…

> When you pay for a Claude subscription, what exactly were you promised?

I was promised 5x or 20x the amount of resources that the free tier would offer. I implicitly expected the same quality too, not some watered-down version of the product they allowed me to sample before committing to a subscription.

Sooner or later Anthropic will run out of VC money, yes. That's their problem, not mine. When I took an Uber while it was subsidized by venture capital, the driver did not drop me half way through my destination because they were having cash flow issues.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#646
post #645

Earlier quoted context omitted.

For context, my company gives each developer a decent monthly allowance for Claude and if push comes to shove, we are allowed to fallback to using AWS Bedrock hosted Anthropic models. When you pay for a Claude subscription, what exactly were you promised? > they will start voting with their money. And go where? Sooner or later the party is going to be over and Claude and its competitors are going to have to start cha…

> When you pay for a Claude subscription, what exactly were you promised? I was promised 5x or 20x the amount of resources that the free tier would offer. I implicitly expected the same quality too, not some watered-down version of the product they allowed me to sample before committing to a subscription. Sooner or later Anthropic will run out of VC money, yes. That's their problem, not mine. When I took an Uber whil…

So how do you know that the free tier hasn’t been reduced by 5x?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#647

Earlier quoted context omitted.

"enshittification" gets thrown around a lot, but this is the exact playbook. Look at the previous bubble's cash cow: advertising. Online advertising is now ubiquitous, terrible, and mandatory for anyone who wants to do e-commerce. You can't run a mass-market online business without buying Adwords, Instagram Ads, etc. AI will be ubiquitous, and then it will get worse and more expensive. But we will be unable to return…

The odds of that happening are high. Trillions invested. It occurred to me an outright rejections of these tools is brewing but can't quite materialise yet.

Trillions promised.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#648
post #339

Earlier quoted context omitted.

The very best open models are maybe 3-12 months behind the frontier and are large enough that you need $10k+ of hardware to run them, and a lot more to run them performantly. ROI here is going to be deeply negative vs just using the same models via API or subscription. You can run smaller models on much more modest hardware but they aren't yet useful for anything more than trivial coding tasks. Performance also reall…

You can also run these models on the cloud with Ollama. You might say what's the difference, but these are models whose performance will stay consistent over time, whether run locally or in the cloud. For $200 a year I'm getting some pretty fantastic results running GLM 5.1 and even Minimax 2.7 and Kimi 2.5 and Gemma 4 on Ollama's cloud instances. And if you don't like Ollama's cloud instance, you can run it on your…

Interesting. On the pricing page, there are still limits placed on the usage. How restrictive have you found them?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#649
Cancelled today after responses became code soup, skills ignored completely, and in response to a question told me "its A, no thats wrong, its B, no actually i dont know, please look for the answer".

Something materially changed in last 4 weeks.

Also, see made up boosterism about finding security holes everywhere. Its just fanning the flames of the industry worries about all the stupid account take overs.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#650

Earlier quoted context omitted.

We are taking it seriously, and are continuing to investigate. We are not trusting the metrics.

The quantitative ux research team at Google was created for exactly this problem: a service which became popular before the right metrics existed, meaning metrics need to be derived first, then optimized. We would observe users (irl), read their logs, then generate experiments to improve the behavior as measured by logs, and return to see if the experiment improves irl experiences. There were not many of us and we ar…

Metrics and quantitative ux results in really bad software, making it rigid while optimizing for the wrong things.

The most obvious example is Google creating multiple steps for Login where you have to enter your password after you put in your user.

I wonder what metric lead to that decision or was it a political decision to make it seem like their "old" software has some new feature.

Post reply on HN