Earlier quoted context omitted.
Can you clearly state what they messed up?
Suddenly burning up the quota ~4x faster than usual is not a mess up in your opinion?
Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
551–560 of 695 posts
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#552Earlier quoted context omitted.
> why would you start with telling people that weren’t able to get an SLA That hasn’t been established. There’s no evidence that they went to Anthropic and tried to negotiate one. > that Anthropic offers SLAs I didn’t. I said “they probably will for the right price .” There are two modifiers in that statement. And the price is unspecified. Their first offer could be a billion dollars. Too expensive? Negotiate down.
I would invite you to notice your interlocutor's assumptions, especially as revealed in his prior comment. Look at how he misunderstands the situation: > If you wanted to convince everybody about a vast universe of secret business and your expertise in it... > Like if I wanted to convince people that In’N’Out has a secret menu... You are discussing business. He is understanding you to be attempting to "mog" him, beca…
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#553I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…
> I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. Indeed. Anthropic is just leading the pack switching to juicy corporate users who are happy to pay thousands per month per dev and leave the fans behind. And now OpenAI is following suit. They lowered significantly the limi…
Granted the experience can be worse, esp. if you're using it very hands-off and not like a junior assistant who's extremely fast but doesn't know what he's doing at the architecture and strategy level. But even for that I'm relatively confident the Chinese will be competitive pretty soon, and they won't be too expensive. And we know this because we can see their current models and we know what it takes to run them.
Currently my Strix Halo computer that costed me under £3k can do a lot of LLM stuff that is perfectly useful. In some ways, it's better than "cloud" models, I have models that essentially don't say "no" and I have relatively predictable setups. If you want to get fancy, you can right now rent compute to run models that are extremely capable like the latest ones from Kimi, GLM, Qwen, Minimax at full size from providers that are not operating at a loss and it won't be too expensive. You can pool resources to do the same locally. You can do stuff that cloud providers are unlikely to market, like distillation and abliteration to serve your specific needs.
I'm very optimistic about open weights models just the way they are right now.
But I agree with you that OpenAI will likely play similar games to Anthropic and it could be soon.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#554Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
I don't want a nudge. I want a clear RED WARNING with "You've gone away from your computer a bit too long and chatted too much at the coffee machine. You're better off starting a new context!"
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#555Earlier quoted context omitted.
I had a weird experience at work last week where Claude was just thinking forever about tasks and not actually doing anything. It was unusable. The next day it was fine again.
That happens to me all the time. My current working theory is when their servers are hammered there is a queueing system that invisible to end-users.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#556Are there local models dedicated to programming already any good? That could be a way to deal with anthropic or others flipflopping with token usage or limits
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#557Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#558Pretty sure OpenCode is not subsidizing, and across Codex 5.x always on xhigh, Claude Opus 4.6 on high effort and a bunch of Chinese models, I only burned about $50 over the last month. I don’t understand why people insist on these subscriptions and CC. Fanboyism is a bit too hardcore at this point. Apple fanboys look extremely prudent compared to this behavior.
For reference, users on claude max 20x who hit their weekly quota would have spent roughly ~$6,000/month in the API. (Source: my own usage) So you just aren't in the same realm of usage. Maybe that is why you don't understand?
What I don’t understand is why people aren’t trying models that are 10x and in some cases 100x cheaper.
Though unclear why you’d assume all my usage would be on Claude Opus when I mentioned “a bunch of Chinese models?”
Unless this is a flex about how many tokens you burned. In which case, congrats...?
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#559Earlier quoted context omitted.
> why would you start with telling people that weren’t able to get an SLA That hasn’t been established. There’s no evidence that they went to Anthropic and tried to negotiate one. > that Anthropic offers SLAs I didn’t. I said “they probably will for the right price .” There are two modifiers in that statement. And the price is unspecified. Their first offer could be a billion dollars. Too expensive? Negotiate down.
I would invite you to notice your interlocutor's assumptions, especially as revealed in his prior comment. Look at how he misunderstands the situation: > If you wanted to convince everybody about a vast universe of secret business and your expertise in it... > Like if I wanted to convince people that In’N’Out has a secret menu... You are discussing business. He is understanding you to be attempting to "mog" him, beca…
I looked up “mogging” and I’d think “my assumptions about stuff are valid because I’m a lawyer and don’t know what you do” would count more as mogging than “that doesn’t quite sound right, this is a conversation about something specific and not your general cleverness” but I’ve got a Benny Hill archive to get through
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#560Earlier quoted context omitted.
The quantitative ux research team at Google was created for exactly this problem: a service which became popular before the right metrics existed, meaning metrics need to be derived first, then optimized. We would observe users (irl), read their logs, then generate experiments to improve the behavior as measured by logs, and return to see if the experiment improves irl experiences. There were not many of us and we ar…
I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.