Earlier quoted context omitted.
Suddenly burning up the quota ~4x faster than usual is not a mess up in your opinion?
It is not inherently their fault though because usage is controlled both by the user and the harness behavior. So I was asking specifically what about the harness was messed up, can you provide that info?
Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
581–590 of 695 posts
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#582Earlier quoted context omitted.
> The value is all the things I built with it? To be clear, the people I were talking about were not referring to the value, but the experience in using these tools. > I also don't understand the "pro-AI" phrase. Would you prefer the phrase "AI-boosters"?
Ah, okay, I must have gotten lost in the conversation. Sorry! > Would you prefer the phrase "AI-boosters"? AI-booster folk? :)
Could be; I mean, we differentiate between people who use cars as a tool and call enthusiasts "petrol-heads".
I use AI daily, but I certainly wouldn't consider myself either pro-AI or an AI-booster.
(Naming is hard)
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#583Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
[flagged]
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#584I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…
IMO we are currently in the ENIAC era of LLMs. Perhaps there will be a brief moment where things get worse, but long term the cost of these things will go way down.
Cumulative AI capex will hit $2T this year. Cumulative opex is on the same order. Unless the models get real good (as in: can fully replace many engineers) right quick, nobody is even going to see interest getting paid on those investments. The only alternative is model access costing 5 figures per (replaced) seat.
But yes, once GPU racks can be had at auction for pennies on the dollar, inference of open source models might be an... OK low margin commodity business.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#585Earlier quoted context omitted.
I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.
[flagged]
They don't use Claude Code, they get accused that they don't even trust it themselves.
They use Claude Code, they get accused the code is shit because it's slop.
I think dogfooding is known to be a legitimate approach here.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#586Earlier quoted context omitted.
[flagged]
This comment seems unnecessarily hostile.
I'm sorry if you and others are offended. They've had these issues for several weeks now. I haven't seen any real improvements during this time. I see more features and more bugs.
There have been several releases made over the last few days without any changelogs. The quotas are still as opaque as they've been. This company has some extremely shady business practices.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#587Earlier quoted context omitted.
I think the parent is saying that one should be aware that the whole LLM industry is still in an experimental stage and far from mature. What you want isn’t what’s being offered. I agree that there should be higher standards, but what we currently have is an arms race. The consequence is to factor that into the value proposition and maybe not rely too much on it.
SLAs should be standard for any paid service, especially on the enterprise side, but also on the consumer side. Being immature as a company does not excuse a lack of service delivery.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#588Earlier quoted context omitted.
[flagged]
This comment seems unnecessarily hostile.
It seems just fine to me. This is what Anthropic needs to do if they want to survive. I'm always looking out for someone to integrate an actually good harness to a good model. Once that happens, I'm jumping ship if Anthropic keeps playing these tricks.
It's almost unusable for me now. A simple prompt to merge 3 sub-100-line files with simple node code, on Sonnet 4.6, uses up 20% of my 5 hour quota, on a new/fresh session.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#589Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
There is only so much you can do through "UX improvements" or some smart routing on the backend. Your flagship product is actively getting worse, and if users need to fiddle with hidden settings and keep track of GitHub issues every week they will start voting with their money.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#590Earlier quoted context omitted.
Boring corporate Ai will surely come, but hey, lets enjoy the wild west while it lasts. I am grateful to see Boris come here to address problems people face. I 100% sure nobody is making him - he has one of the coolest jobs in the world.
>he has one of the coolest jobs in the world. So that means we just eject any critical thinking when it comes to companies, especially where they is no liability or obligation for them (Boris or Anthropic) to be honest. Other than 'trust'.