Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

621–630 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#621
post #91

Earlier quoted context omitted.

Lights on = Ads in your output. EOY latest; they can't keep kicking the massive costs down the road.

Where is your evidence of this "massive cost"? Inference is massively profitable for both anthropic and openai. Training is not.

Inference cannot happen without training the model first, so the distinction is quite pointless.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#622
post #589

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

So Anthropic is trying to save money on infrastructure, we all get it. However, it's not ok to degrade the performance your users have paid for. Last week the issue was that you reduced the default "effort" level, now the prompt cache is shortened. Several users experience far more restrictive usage limits lately. There is only so much you can do through "UX improvements" or some smart routing on the backend. Your fl…

Where did they say the prompt cache is shortened?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#623

Earlier quoted context omitted.

Anthropic can't win in this case. They don't use Claude Code, they get accused that they don't even trust it themselves. They use Claude Code, they get accused the code is shit because it's slop. I think dogfooding is known to be a legitimate approach here.

The idea is that Claude Code is surprisingly buggy and unrefined for something created by the very tool and processes that are supposed to be replacing us as we speak.

The idea is that sculpted ideal code is rarely the best choice.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#624

Earlier quoted context omitted.

How can we turn of 1m context? I don't find it has ever helped.

He mentioned this in his original comment: "CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000"

There's also CLAUDE_CODE_DISABLE_1M_CONTEXT and I'm really not clear on what the difference is and why to pick one over the other. But I guess one disables models that have 1m and the other keeps those models but sets the limit lower?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#625
post #589

Earlier quoted context omitted.

So Anthropic is trying to save money on infrastructure, we all get it. However, it's not ok to degrade the performance your users have paid for. Last week the issue was that you reduced the default "effort" level, now the prompt cache is shortened. Several users experience far more restrictive usage limits lately. There is only so much you can do through "UX improvements" or some smart routing on the backend. Your fl…

Where did they say the prompt cache is shortened?

from 1h to 5 minutes, was in the news recently

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#626
post #546

Earlier quoted context omitted.

I am sorry you feel this way, but the reality of the situation is there is zero reason to trust anything Anthropic or Boris says. They have no legal liability or obligation to tell the truth, besides brand risk, which to people like you is mitigated for a single person to show up, post, and thats it.

You should work at these companies and understand they have good intentioned employees otherwise they’d rarely pass the cultural interviews plus background checks plus backchanneling. Have a bit more faith in the employees

> Have a bit more faith in the employees

Have you been asleep for a decade?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#627

Earlier quoted context omitted.

[flagged]

Anthropic can't win in this case. They don't use Claude Code, they get accused that they don't even trust it themselves. They use Claude Code, they get accused the code is shit because it's slop. I think dogfooding is known to be a legitimate approach here.

And they don't use our version of CC, or with our settings. They have flags for internal use only.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#628

Earlier quoted context omitted.

SLAs should be standard for any paid service, especially on the enterprise side, but also on the consumer side. Being immature as a company does not excuse a lack of service delivery.

Not every customer, even a paying customer, demands reliability at a particular level. Market segmentation tends to address those situations: pay more, get more.

> pay more, get more

Users on $200 plan complaining, already at max level of subscription, I don't think a $200 subscription should make you feel like you are getting unfair advantage. Like restricting claude -p to API ... after I paid so much? Moderate use should not do that. I am not running it batch mode on a million inputs.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#629
post #339
post #90

How good are local LLMs at coding these days? Does anyone have any recommendations for how to get this setup? What would the minimum spend be for usable hardware? I am getting bored of having to plan my weekends around quota limit reset times...

The very best open models are maybe 3-12 months behind the frontier and are large enough that you need $10k+ of hardware to run them, and a lot more to run them performantly. ROI here is going to be deeply negative vs just using the same models via API or subscription. You can run smaller models on much more modest hardware but they aren't yet useful for anything more than trivial coding tasks. Performance also reall…

You can also run these models on the cloud with Ollama. You might say what's the difference, but these are models whose performance will stay consistent over time, whether run locally or in the cloud. For $200 a year I'm getting some pretty fantastic results running GLM 5.1 and even Minimax 2.7 and Kimi 2.5 and Gemma 4 on Ollama's cloud instances. And if you don't like Ollama's cloud instance, you can run it on your own cloud instance from the very same providers that Ollama is using. They use NVIDIA cloud providers (NCPs) although not sure which ones specifically and claims that the "cloud does not retain your data to ensure privacy and security." [https://ollama.com/blog/cloud-models]

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#630

Earlier quoted context omitted.

"enshittification" gets thrown around a lot, but this is the exact playbook. Look at the previous bubble's cash cow: advertising. Online advertising is now ubiquitous, terrible, and mandatory for anyone who wants to do e-commerce. You can't run a mass-market online business without buying Adwords, Instagram Ads, etc. AI will be ubiquitous, and then it will get worse and more expensive. But we will be unable to return…

But why would they make the product shittier and not just more expensive? A lot of the complaints have been the model getting lost and going rogue.

Why isn't there a premium, ad-free Google Search (or Facebook, or Instagram)? Because the most valuable customers (with the most money) self-select out of seeing ads. It would collapse the 2-sided market and create a race to the bottom. There is a dollar amount of advertising revenue per customer, but as John Wannamaker said - "Half the money I spend on advertising is wasted, but I don't know which half".

If the AI companies made their pricing "pay as you go" without quotas, a few insane zealots (power users) would occupy all the capacity and choke everyone else out. Regardless of the cost, the AI providers would lose the ubiquity they currently enjoy, and become a niche tool for rich tech people. They would rather be a mile wide and an inch deep, doing a worse job serving millions of users, because there's a better scaling narrative for legislating and fundraising that way. Like the advertisers there are intolerable indirect effects of letting valuable "power users" spend more money to get a better experience.

Post reply on HN