Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

671–680 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#671

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

One thing I didn't see anywhere here, except your mention about pulling in large number of skills, is that the token consumption is significantly higher for users with many agents, skills, and MCPs installed, and many are mere ghosts. The 5m TTL from #46829 compounds the effect: in my case, I found ~20k tokens of ghost context I hadn't intentionally opened. Each idle period after 5m wastes that as a full cache miss.

Boris, would you please confirm on-record: is the current cache TTL for the main agent context 1h or 5m? Issue #46829 was closed as "not planned".

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#672
post #474

Earlier quoted context omitted.

I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.

But no other user has yet come and said "I worked with ajma in the past ..." so how can we trust your judgement about Boris?

I saw this guy named Claude saying ajma is a genius!

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#673
post #474

Earlier quoted context omitted.

The quantitative ux research team at Google was created for exactly this problem: a service which became popular before the right metrics existed, meaning metrics need to be derived first, then optimized. We would observe users (irl), read their logs, then generate experiments to improve the behavior as measured by logs, and return to see if the experiment improves irl experiences. There were not many of us and we ar…

I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.

Nice try boris

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#674

Earlier quoted context omitted.

Metrics and quantitative ux results in really bad software, making it rigid while optimizing for the wrong things. The most obvious example is Google creating multiple steps for Login where you have to enter your password after you put in your user. I wonder what metric lead to that decision or was it a political decision to make it seem like their "old" software has some new feature.

If you mean Google website login, that step is needed because the email address is used to determine which identity provider to use. E.g. I have three different accounts that branch off from that same initial login flow. One is my person "gmail.com" account, and the other two go through enteprise identity providers related to my employment and their G-Suite licenses. So after I put in one of these three email address…

I mean I get it logically makes sense. But it still seems like a waste of time for a small percentage of use cases.

Maybe a better approach is put in your login have it automatically detect if it requires an identity provider. Gray out the password to signal to the user password is not necessary and automatically redirect.

Less clicking, don't break flow and think of a smoother solution.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#675

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

you should check with people working on Claude Code, cache has been udpated to 5min ... https://github.com/anthropics/claude-code/issues/46829#issue...

So yeah, 1M window that expires every 5min .... not good

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#676
post #483

I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…

> I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. Indeed. Anthropic is just leading the pack switching to juicy corporate users who are happy to pay thousands per month per dev and leave the fans behind. And now OpenAI is following suit. They lowered significantly the limi…

[dead]

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#677

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

it seems if context can't be held for over an hour it should warn you a countdown or such; i already enabled the tokens verbosity thing to see what token level i'm at, but i often leave things sitting rather than complete so that i'm tying things up to start something new in the morning rather than starting on a new thing. so like i just resumed a session that was near-complete, and now it's gone and reloaded all that session in? bit i hadn't detached it. i kind of thougth /summary itself had to read the whole token flow, but that the token context was held locally for some reason..

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#679

Earlier quoted context omitted.

> You've embarrassed yourself quite badly, I'm afraid. :( you are right. This isn’t the first time I’ve lost an argument because hours into a discussion somebody introduced “what if a billion dollars” or “magic amulet” or “ブルマの母” etc

A billion dollars is just an example. I could have said a million. When someone says "a high price" that's unspecified, you can use your imagination to hazard a guess at what that might be. Such a figure might seem unreasonable or unrealistic to you, but deals are done between companies under terms most individuals wouldn't come close to considering. The only reason I mentioned being an attorney was because someone i…

https://i.kym-cdn.com/entries/icons/facebook/000/034/711/Scr...

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#680

Earlier quoted context omitted.

I appreciate your kindness. While I’ve got you, did you know that the Benny Hill show started in 1955 and a good chunk of what aired from then to 1969 was lost? There are a lot of fans that don’t even realize that what is sometimes labeled as season 1 is season 15! Crazy stuff!

I had not known that! In a similar vein, there exists an Alice in Wonderland -themed Muppet Show episode, starring Brooke Shields, which has had to be left out of home video releases due to so far unresolvable music licensing issues. Not quite totally lost, but somewhat hard to find!

I’ll check that out! If I find a good link for it I’ll post it as a reply here.
Post reply on HN