Live data from Hacker News

Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

github.com

551–560 of 695 posts

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#552

Earlier quoted context omitted.

> why would you start with telling people that weren’t able to get an SLA That hasn’t been established. There’s no evidence that they went to Anthropic and tried to negotiate one. > that Anthropic offers SLAs I didn’t. I said “they probably will for the right price .” There are two modifiers in that statement. And the price is unspecified. Their first offer could be a billion dollars. Too expensive? Negotiate down.

I would invite you to notice your interlocutor's assumptions, especially as revealed in his prior comment. Look at how he misunderstands the situation: > If you wanted to convince everybody about a vast universe of secret business and your expertise in it... > Like if I wanted to convince people that In’N’Out has a secret menu... You are discussing business. He is understanding you to be attempting to "mog" him, beca…

[deleted]

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#553
post #483

I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. For those not in the Google Gemini/Antigravity sphere, over the last month or so that community has been experiencing nothing short of contempt from Google when attempting to address an apparent bait and switch on quota expec…

> I'm afraid the music may be slowly fading at this party, and the lights will soon be turned on. We may very well look back on the last couple years as the golden era of subsidized GenAI compute. Indeed. Anthropic is just leading the pack switching to juicy corporate users who are happy to pay thousands per month per dev and leave the fans behind. And now OpenAI is following suit. They lowered significantly the limi…

I don't know about users on reddit and discord, but the open models are essentially at SotA with a 3-4 months delay. That puts a hard backstop at what OpenAI and Anthropic can do before I personally can cut them off entirely without losing too much.

Granted the experience can be worse, esp. if you're using it very hands-off and not like a junior assistant who's extremely fast but doesn't know what he's doing at the architecture and strategy level. But even for that I'm relatively confident the Chinese will be competitive pretty soon, and they won't be too expensive. And we know this because we can see their current models and we know what it takes to run them.

Currently my Strix Halo computer that costed me under £3k can do a lot of LLM stuff that is perfectly useful. In some ways, it's better than "cloud" models, I have models that essentially don't say "no" and I have relatively predictable setups. If you want to get fancy, you can right now rent compute to run models that are extremely capable like the latest ones from Kimi, GLM, Qwen, Minimax at full size from providers that are not operating at a loss and it won't be too expensive. You can pool resources to do the same locally. You can do stuff that cloud providers are unlikely to market, like distillation and abliteration to serve your specific needs.

I'm very optimistic about open weights models just the way they are right now.

But I agree with you that OpenAI will likely play similar games to Anthropic and it could be soon.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#554

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

I don't want a nudge. I want a clear RED WARNING with "You've gone away from your computer a bit too long and chatted too much at the coffee machine. You're better off starting a new context!"

forget the warning, just compact like someone suggested in the ticket. Who would opt for a massive cache miss?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#555
post #115
post #75

Earlier quoted context omitted.

I had a weird experience at work last week where Claude was just thinking forever about tasks and not actually doing anything. It was unusable. The next day it was fine again.

That happens to me all the time. My current working theory is when their servers are hammered there is a queueing system that invisible to end-users.

The way Claude/Codex behave is entirely consistent with how every vibe coded project (of mine) has ended up so far. I bet those guys have no idea what's going on and are taking guesses because no one understands the thing they've made.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#556

Are there local models dedicated to programming already any good? That could be a way to deal with anthropic or others flipflopping with token usage or limits

Short answer no. Less short answer, the science is catching up to big ones quickly.

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#557

Hey all, Boris from the Claude Code team here. We've been investigating these reports, and a few of the top issues we've found are: 1. Prompt cache misses when using 1M token context window are expensive. Since Claude Code uses a 1 hour prompt cache window for the main agent, if you leave your computer for over an hour then continue a stale session, it's often a full cache miss. To improve this, we have shipped a few…

Number 2 makes me chuckle honestly. Too many people going down the 10x rabbit holes on youtube. Next up, a framework that 100xs your workflow. You know its good because it comes with 300 agents and 20 mcp servers and 1200 skills

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#558
post #365

Pretty sure OpenCode is not subsidizing, and across Codex 5.x always on xhigh, Claude Opus 4.6 on high effort and a bunch of Chinese models, I only burned about $50 over the last month. I don’t understand why people insist on these subscriptions and CC. Fanboyism is a bit too hardcore at this point. Apple fanboys look extremely prudent compared to this behavior.

For reference, users on claude max 20x who hit their weekly quota would have spent roughly ~$6,000/month in the API. (Source: my own usage) So you just aren't in the same realm of usage. Maybe that is why you don't understand?

I guess I could’ve been clearer.

What I don’t understand is why people aren’t trying models that are 10x and in some cases 100x cheaper.

Though unclear why you’d assume all my usage would be on Claude Opus when I mentioned “a bunch of Chinese models?”

Unless this is a flex about how many tokens you burned. In which case, congrats...?

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#559

Earlier quoted context omitted.

> why would you start with telling people that weren’t able to get an SLA That hasn’t been established. There’s no evidence that they went to Anthropic and tried to negotiate one. > that Anthropic offers SLAs I didn’t. I said “they probably will for the right price .” There are two modifiers in that statement. And the price is unspecified. Their first offer could be a billion dollars. Too expensive? Negotiate down.

I would invite you to notice your interlocutor's assumptions, especially as revealed in his prior comment. Look at how he misunderstands the situation: > If you wanted to convince everybody about a vast universe of secret business and your expertise in it... > Like if I wanted to convince people that In’N’Out has a secret menu... You are discussing business. He is understanding you to be attempting to "mog" him, beca…

I am so old :(

I looked up “mogging” and I’d think “my assumptions about stuff are valid because I’m a lawyer and don’t know what you do” would count more as mogging than “that doesn’t quite sound right, this is a conversation about something specific and not your general cleverness” but I’ve got a Benny Hill archive to get through

Re: Pro Max 5x quota exhausted in 1.5 hours despite moderate usage

#560
post #474

Earlier quoted context omitted.

The quantitative ux research team at Google was created for exactly this problem: a service which became popular before the right metrics existed, meaning metrics need to be derived first, then optimized. We would observe users (irl), read their logs, then generate experiments to improve the behavior as measured by logs, and return to see if the experiment improves irl experiences. There were not many of us and we ar…

I worked with Boris in the past and in my experience, Boris cares deeply about the customer. I'd vouch that Boris really cares about the issue people are running into.

[flagged]
Post reply on HN