Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

581–590 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#581

Earlier quoted context omitted.

Suddenly burning up the quota ~4x faster than usual is not a mess up in your opinion?

It is not inherently their fault though because usage is controlled both by the user and the harness behavior. So I was asking specifically what about the harness was messed up, can you provide that info?

It's all there, including the specific version regression, unearthed bugs, workarounds: https://github.com/anthropics/claude-code/issues/45756

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#582

Earlier quoted context omitted.

> The value is all the things I built with it? To be clear, the people I were talking about were not referring to the value, but the experience in using these tools. > I also don't understand the "pro-AI" phrase. Would you prefer the phrase "AI-boosters"?

Ah, okay, I must have gotten lost in the conversation. Sorry! > Would you prefer the phrase "AI-boosters"? AI-booster folk? :)

> AI-booster folk? :)

Could be; I mean, we differentiate between people who use cars as a tool and call enthusiasts "petrol-heads".

I use AI daily, but I certainly wouldn't consider myself either pro-AI or an AI-booster.

(Naming is hard)

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#583

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

[flagged]

This comment seems unnecessarily hostile.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#584

I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…

IMO we are currently in the ENIAC era of LLMs. Perhaps there will be a brief moment where things get worse, but long term the cost of these things will go way down.

I assume the "briefly gets worse" is when a buch of hyperscalers do a complete write-off of their entire AI investments, bankrupting several of them (which, in turn, bankrupts several large banks and most current venture capital firms)?

Cumulative AI capex will hit $2T this year. Cumulative opex is on the same order. Unless the models get real good (as in: can fully replace many engineers) right quick, nobody is even going to see interest getting paid on those investments. The only alternative is model access costing 5 figures per (replaced) seat.

But yes, once GPU racks can be had at auction for pennies on the dollar, inference of open source models might be an... OK low margin commodity business.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#585
post #474

Earlier quoted context omitted.

I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.

[flagged]

Anthropic can't win in this case.

They don't use Claude Code, they get accused that they don't even trust it themselves.

They use Claude Code, they get accused the code is shit because it's slop.

I think dogfooding is known to be a legitimate approach here.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#586

Earlier quoted context omitted.

[flagged]

This comment seems unnecessarily hostile.

They appear to take issues seriously mostly when they become posts on hacker news and when articles are published online by major news sites. Customer support is mostly a bot. I don't even know how to reach some actual humans to get support.

I'm sorry if you and others are offended. They've had these issues for several weeks now. I haven't seen any real improvements during this time. I see more features and more bugs.

There have been several releases made over the last few days without any changelogs. The quotas are still as opaque as they've been. This company has some extremely shady business practices.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#587
post #402

Earlier quoted context omitted.

I think the parent is saying that one should be aware that the whole LLM industry is still in an experimental stage and far from mature. What you want isn’t what’s being offered. I agree that there should be higher standards, but what we currently have is an arms race. The consequence is to factor that into the value proposition and maybe not rely too much on it.

SLAs should be standard for any paid service, especially on the enterprise side, but also on the consumer side. Being immature as a company does not excuse a lack of service delivery.

If you can point to a consumer targeted service that provides and keeps their SLAs, I’ll be impressed.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#588

Earlier quoted context omitted.

[flagged]

This comment seems unnecessarily hostile.

Why?

It seems just fine to me. This is what Anthropic needs to do if they want to survive. I'm always looking out for someone to integrate an actually good harness to a good model. Once that happens, I'm jumping ship if Anthropic keeps playing these tricks.

It's almost unusable for me now. A simple prompt to merge 3 sub-100-line files with simple node code, on Sonnet 4.6, uses up 20% of my 5 hour quota, on a new/fresh session.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#589

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

So Anthropic is trying to save money on infrastructure, we all get it. However, it's not ok to degrade the performance your users have paid for. Last week the issue was that you reduced the default "effort" level, now the prompt cache is shortened. Several users experience far more restrictive usage limits lately.

There is only so much you can do through "UX improvements" or some smart routing on the backend. Your flagship product is actively getting worse, and if users need to fiddle with hidden settings and keep track of GitHub issues every week they will start voting with their money.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#590
post #364

Earlier quoted context omitted.

Boring corporate Ai will surely come, but hey, lets enjoy the wild west while it lasts. I am grateful to see Boris come here to address problems people face. I 100% sure nobody is making him - he has one of the coolest jobs in the world.

>he has one of the coolest jobs in the world. So that means we just eject any critical thinking when it comes to companies, especially where they is no liability or obligation for them (Boris or Anthropic) to be honest. Other than 'trust'.

Don’t like Anthropic? Use a competing service. At this point the sheer volume of your commentary is not particularly complimentary to your own critical thinking skills. It’s not your job to correct the internet or to convince randoms of the rightness of your position. Of all the things in the world to be pissed at so insistently, this seems to be a pretty minor one.
Post reply on HN