Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

571–580 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#571

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

Claude Code cache is not 1 hour. There is a "Closed as not planned" issue in GitHub that confirms that it has been moved to 5 minutes since March: https://github.com/anthropics/claude-code/issues/46829 . I started seeing the massive degradation exactly on the 23rd of March, hence after a few days I unsubscribed because it was completely unusable, with a ~5h session being depleted in as little as 15-20 mins.

Looks like the cache change to 5 minutes was so secretive that even CC team doesn't know that.

Or someone just vibe coded "Hey, Claude, make them burn allowances quicker" and merged without telling anyone.

Both are plausible to me.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#572

I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…

Maybe I missed the party, but it feels like it's just starting. I have only been running local models and we are finally at the point with gemma4 and Qwen3.5 where they can start doing coding work. And the quota can't change.

I am surprisingly optimistic about local LLMs. Their progress (especially with regards to distillation) over the last year has been remarkable. Qwen 3.5 is amazing for what it is. It think it's production capable - for many use cases, but not all. It does require more careful alignment of instructions, and offers a smaller context (even with very large unified memory). But with some care, one can code all day, every day, without limits. The Mac Mini 64GB is probably sufficient for Qwen 3.5 35B. Go larger for larger contexts.

Of course it's not as easy as pointing February Opus 4.6 at a folder and giving it one-sentence instructions.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#573

Earlier quoted context omitted.

The value is all the things I built with it? Surely, this constant change deteriorates the experience but to be clear, here we're nitpicking on the experience, not questioning the value. I also don't understand the "pro-AI" phrase. It's a tool, it brings results. I'm not pro-car when I drive to work.

> The value is all the things I built with it? To be clear, the people I were talking about were not referring to the value, but the experience in using these tools. > I also don't understand the "pro-AI" phrase. Would you prefer the phrase "AI-boosters"?

Ah, okay, I must have gotten lost in the conversation. Sorry!

> Would you prefer the phrase "AI-boosters"?

AI-booster folk? :)

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#574

Earlier quoted context omitted.

> You've embarrassed yourself quite badly, I'm afraid. :( you are right. This isn’t the first time I’ve lost an argument because hours into a discussion somebody introduced “what if a billion dollars” or “magic amulet” or “ブルマの母” etc

It's just a world you've never seen. Don't take it too personally.

I appreciate your kindness. While I’ve got you, did you know that the Benny Hill show started in 1955 and a good chunk of what aired from then to 1969 was lost? There are a lot of fans that don’t even realize that what is sometimes labeled as season 1 is season 15! Crazy stuff!

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#575
post #24

Unless the agent code is open-sourced, there is hardly any transparency in how the agent is spending your tokens and how does it calculate the tokens. It's like asking your lawyer why they charged some amount.

You can insert a proxy in between and look at precisely what it is sending if you’re so inclined

CC accepts http endpoints so doesn’t require anything too complicated

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#576

I’ve switched to open code and openrouter. I only did the $20/month subscription since 9/2025 It was great for about 5 months, amazing in fact. I under utilized it. For the past month, it’s basically unusable, both Claude code and just Claude chat. 1-2 prompts and I’m out. Last week I prob sent a total of 15 messages to Claude and was out of daily and weekly usage each day. I get that the $20/month subscription isn’t…

I'm curious if people are going to be switching to something else. OpenAI perhaps?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#577

Earlier quoted context omitted.

Dang man, chill.

Man, expecting the minimal from companies who are supposed to deliver a pro... there is no SLA for any this, so you are right. Also, why is there no SLA?

You’re not getting a worthwhile sla on a subscription at this rate. What are you going to get? A few dollars? An sla isn’t useful unless it actually bites for the provider and actually compensates the customer. And it costs money - how much are you willing to spend for this insurance?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#578
post #332

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

Hello Boris! How do I increase the 1 hour prompt cache window for the main agent? I would love to be able to set that to, say, 4 hours. That gives me enough time to work on something, go teach a class, grab a snack, and come back and pick up where I left off.

Another CC team member confirmed it's 5 minutes now, not 1 hour.

See the links in https://news.ycombinator.com/item?id=47747209

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#579

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

[flagged]

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#580

Earlier quoted context omitted.

Suddenly burning up the quota ~4x faster than usual is not a mess up in your opinion?

[flagged]

LOL, funny how you're so happy to dismiss dozens of reports with hard data, and confirmed by the Claude Code team member.

Issue with the confirmation: https://github.com/anthropics/claude-code/issues/45756

Looks like you have an axe to grind and facts be damned? :D

Post reply on HN