Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

411–420 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#411
post #294

Earlier quoted context omitted.

Can’t you turn the 1M off with a /model opus (or /model sonnet)? At least up until recently the 1M model was separated into /model opus[1M]

1M context window is still a separate, non-default model in Claude Code and not included with subscriptions (billed at API rates only)

what? Opus 1m has been in place for at least a few weeks for plan users.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#412

I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…

Ultimately we'll find more efficient techniques and hardware and AI companies will end up owning Nuclear Power Stations and continue providing models capable of 10x of what they are now. Valuation have already reached point where these companies can run their nuclear power station, fund developement of new hardware and techniques and boost capabilities of their models by 10x

There's not enough nuclear to go around, and the approval/permitting process for new nuclear power plants is nothing to sneeze at, both in terms of time and cost.

That's also ignoring that nuclear power plants also consume quite a bit of water, which may be a more difficult bottleneck in and of itself even without trying to add nuclear into the mix.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#413
post #402

Earlier quoted context omitted.

Wait, where is there a 'beta' tag to something that they are charging real money for? Why is this software any different than any other software and we should completely give away our rights as a consumer to ensure what we pay for is delivered?

I think the parent is saying that one should be aware that the whole LLM industry is still in an experimental stage and far from mature. What you want isn’t what’s being offered. I agree that there should be higher standards, but what we currently have is an arms race. The consequence is to factor that into the value proposition and maybe not rely too much on it.

SLAs should be standard for any paid service, especially on the enterprise side, but also on the consumer side. Being immature as a company does not excuse a lack of service delivery.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#414

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

[deleted]

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#415

Earlier quoted context omitted.

What you say makes sense, but they very actually reduced the token limits. We had, say, 20M tokens/week before, now we have 18M tokens/week (example numbers). They didn't just make a model that eats tokens faster.

Is that documented somewhere? Do you have a link? I might have just missed it, and if they did it, I will take my words back.

https://www.ghacks.net/2026/03/27/anthropic-reduces-claude-s...

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#416

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

There's an issue someone raised showing that prompt caches are only 5 minutes.

The reply seems to be: oh huh, interesting. Maybe that's a good thing since people sometimes one-shot? That doesn't feel like the messaging I want to be reading, and the way it conflicts with the message here that cache is 1 hour is confusing.

https://news.ycombinator.com/item?id=47741755

Is there any status information or not on whether cache is used? It sure looks like the person analyzing the 5m issue had to work extremely hard to get any kind of data. It feels like the iteration loop of people getting better at this stuff would go much much better if this weren't such a black box, if we had the data to see & understand: is the cache helping?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#417

Earlier quoted context omitted.

After googling https://www.reddit.com/r/singularity/comments/1psesym/openai...

I've seen sources like this before. It's all hearsay and promo. I was asking for any publicly available verifiable information regarding the cost of inference at scale. I haven't seen any such info personally which is why I asked. I'm dying to see S-1 filing for Anthropic or OpenAI. I don't actually think inference is as cheap as people say if you consider the total cost (hardware, energy, capex, etc)

Well they're not public yet so you'll have to put up with rumors. But the numbers are available for companies like DeepSeek say they have an 80% profit margin, so it stands to reason OAI etc would do similar numbers considering they charge much more.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#418

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

I wish people would pay more attention to: * Anthropic is in some way trying to run a business (not a charity) and at least (eventually?) make money and not subsidize usage forever * "What a steal/good deal" the $100-$200/mo plans are compared to if they had to pay for raw API usage and less on "how dare you reserve the right to tweak the generous usage patterns you open-ended-ly gave us, we are owed something!"

As an (ex) paying customer, I'm expecting some consistency. I used to be satisfied with the value I got, until the limits changed overnight, and I'd get a ten of my previous usage.

If Anthropic is allowed to alter the deal whenever, then I'd expect to be able to get my money back, pro-rata, no questions asked.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#419

Been running into the same issue since a week or 2 ago on Opus. To be fair I have a pretty loose harness and pattern but it’s been enough to pull in 20k in bounties a month for a long time without going over plan with very little steering (sometimes days of continuous work) That being said I’ve figured this was coming for a long time and have been slowly moving to local models. They’re slower but with the right harne…

You're really completing bug bounties with found with AI? are companies honoring these?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#420

Earlier quoted context omitted.

Is that documented somewhere? Do you have a link? I might have just missed it, and if they did it, I will take my words back.

https://www.ghacks.net/2026/03/27/anthropic-reduces-claude-s...

Ah yes this is sad to see and a lame move for sure… It’s indeed dependent on usage hours but it’s a bad move even if I’m personally not affected since I use it outside of those hours but I agree it’s lame….
Post reply on HN