Earlier quoted context omitted.
Can’t you turn the 1M off with a /model opus (or /model sonnet)? At least up until recently the 1M model was separated into /model opus[1M]
1M context window is still a separate, non-default model in Claude Code and not included with subscriptions (billed at API rates only)
Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
411–420 of 695 posts
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#412I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…
Ultimately we'll find more efficient techniques and hardware and AI companies will end up owning Nuclear Power Stations and continue providing models capable of 10x of what they are now. Valuation have already reached point where these companies can run their nuclear power station, fund developement of new hardware and techniques and boost capabilities of their models by 10x
That's also ignoring that nuclear power plants also consume quite a bit of water, which may be a more difficult bottleneck in and of itself even without trying to add nuclear into the mix.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#413Earlier quoted context omitted.
Wait, where is there a 'beta' tag to something that they are charging real money for? Why is this software any different than any other software and we should completely give away our rights as a consumer to ensure what we pay for is delivered?
I think the parent is saying that one should be aware that the whole LLM industry is still in an experimental stage and far from mature. What you want isn’t what’s being offered. I agree that there should be higher standards, but what we currently have is an arms race. The consequence is to factor that into the value proposition and maybe not rely too much on it.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#414Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#415Earlier quoted context omitted.
What you say makes sense, but they very actually reduced the token limits. We had, say, 20M tokens/week before, now we have 18M tokens/week (example numbers). They didn't just make a model that eats tokens faster.
Is that documented somewhere? Do you have a link? I might have just missed it, and if they did it, I will take my words back.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#416Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
The reply seems to be: oh huh, interesting. Maybe that's a good thing since people sometimes one-shot? That doesn't feel like the messaging I want to be reading, and the way it conflicts with the message here that cache is 1 hour is confusing.
https://news.ycombinator.com/item?id=47741755
Is there any status information or not on whether cache is used? It sure looks like the person analyzing the 5m issue had to work extremely hard to get any kind of data. It feels like the iteration loop of people getting better at this stuff would go much much better if this weren't such a black box, if we had the data to see & understand: is the cache helping?
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#417Earlier quoted context omitted.
After googling https://www.reddit.com/r/singularity/comments/1psesym/openai...
I've seen sources like this before. It's all hearsay and promo. I was asking for any publicly available verifiable information regarding the cost of inference at scale. I haven't seen any such info personally which is why I asked. I'm dying to see S-1 filing for Anthropic or OpenAI. I don't actually think inference is as cheap as people say if you consider the total cost (hardware, energy, capex, etc)
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#418Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…
I wish people would pay more attention to: * Anthropic is in some way trying to run a business (not a charity) and at least (eventually?) make money and not subsidize usage forever * "What a steal/good deal" the $100-$200/mo plans are compared to if they had to pay for raw API usage and less on "how dare you reserve the right to tweak the generous usage patterns you open-ended-ly gave us, we are owed something!"
If Anthropic is allowed to alter the deal whenever, then I'd expect to be able to get my money back, pro-rata, no questions asked.
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#419Been running into the same issue since a week or 2 ago on Opus. To be fair I have a pretty loose harness and pattern but it’s been enough to pull in 20k in bounties a month for a long time without going over plan with very little steering (sometimes days of continuous work) That being said I’ve figured this was coming for a long time and have been slowly moving to local models. They’re slower but with the right harne…
Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
#420Earlier quoted context omitted.
Is that documented somewhere? Do you have a link? I might have just missed it, and if they did it, I will take my words back.
https://www.ghacks.net/2026/03/27/anthropic-reduces-claude-s...