This is the reality I'm seeing too. Does this mean that the subscriptions (5x, 10x, 20x) are essentially reduced in token-count by 20-30%?
Measuring Claude 4.7's tokenizer costs
201–210 of 540 posts
Re: Measuring Claude 4.7's tokenizer costs
#202Earlier quoted context omitted.
But the unique thing about AI is that the "world" is depending on it like water, oil, gas, etc. Not just a specific use case.
So it should be free? What's your point exactly?
In this context I also imagine we will have greater and greater local models, and the (dependency) ending game is completely unclear.
Re: Measuring Claude 4.7's tokenizer costs
#203Earlier quoted context omitted.
A reasonable conclusion, considering that money and power seem to have their own gravity, so people with more of both end up getting even more of both, and vice versa. Can't blame someone who comes to such a conclusion about money and power.
The unreasonable part automatically labeling power as evil.
Re: Measuring Claude 4.7's tokenizer costs
#204Earlier quoted context omitted.
A reasonable conclusion, considering that money and power seem to have their own gravity, so people with more of both end up getting even more of both, and vice versa. Can't blame someone who comes to such a conclusion about money and power.
The unreasonable part automatically labeling power as evil.
Re: Measuring Claude 4.7's tokenizer costs
#205Earlier quoted context omitted.
I mean, the signs have been there that the costs to run and operate these models wasn't as simple as inference costs. And the signs were there (and, arguably, are still there) that it costs way, way more than many people like to claim on the part of Anthropic. So to me this price hike is not at all surprising. It was going to come eventually, and I suspect it's nowhere near over. It wouldn't surprise me if in 2-3 yea…
> It wouldn't surprise me if in 2-3 years the "max" plan is $800 or $2000 even. I'd rather be surprised if they are still doing business by then.
I’m guessing we’re gonna have a world like working on cars - most people won’t have expensive tools (ex a full hydraulic lift) for personal stuff, they are gonna have to make do with lesser tools.
Re: Measuring Claude 4.7's tokenizer costs
#206Earlier quoted context omitted.
> It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost. I think we're reaching the point where more developers need to start right-sizing the model and effort level to the task. It was easy to get comfortable with using the best model at the highest setting for every…
Except developers can’t even do that. Estimation of any not-small task that hasn’t been done before is essentially a random guess.
Re: Measuring Claude 4.7's tokenizer costs
#207Earlier quoted context omitted.
I want to give give you realistic expectations: Unless you spend well over $10K on hardware, you will be disappointed, and will spend a lot of time getting there. For sophisticated coding tasks, at least. (For simple agentic work, you can get workable results with a 3090 or two, or even a couple 3060 12GBs for half the price. But they're pretty dumb, and it's a tease. Hobby territory, lots of dicking around.) Do your…
We need more voices like this to cut through the bullshit. It's fine that people want to tinker with local models, but there has been this narrative for too long that you can just buy more ram and run some small to medium sized model and be productive that way. You just can't, a 35b will never perform at the level of the same gen 500b+ model. It just won't and you are basically working with GPT-4 (the very first one…
Open models are not bullshit, they work fine for many cases and newer techniques like SSD offload make even 500B+ models accessible for simple uses (NOT real-time agentic coding!) on very limited hardware. Of course if you want the full-featured experience it's going to cost a lot.
Re: Measuring Claude 4.7's tokenizer costs
#208Just yesterday I was happy to have gotten my weekly limit reset [1]. And although I've been doing a lot of mockup work (so a lot of HTML getting written), I think the 1M token stuff is absolutely eating up tokens like CRAZY. I'm already at 27% of my weekly limit in ONE DAY. https://news.ycombinator.com/item?id=47799256
> I'm already at 27% of my weekly limit in ONE DAY. Ouch, that's very different than experience. What effort level? Are you careful to avoid pushing session context use beyond 350k or so (assuming 1m context)?
And this particular set of things has context routinely hit 350-450k before I compact.
That's likely what it is? I think this particular work stream is eating a lot of tokens.
Earlier this week (before Open 4.7 hit), I just turned off 1m context and had it grow a lot slower.
I also have it on high all the time. Medium was starting to feel like it was making the occasional bad decisions and also forgetting things more.
Re: Measuring Claude 4.7's tokenizer costs
#209Re: Measuring Claude 4.7's tokenizer costs
#210Earlier quoted context omitted.
It's unsurprising when this is the first day that tokens have been crazy like this. All of us doing crazy agentic stuff were fine on max before this. Now with Opus 4.7, we're no longer fine, and troubleshooting, and working through options.
> were fine on max before this Ya...you may be who I'm talking about though (if you're speaking from experience). If your methodology is "I used 4.6 max, so I'm going to try 4.7 max" this is fully on you - 4.7 max is not equivalent to 4.6 max, you want 4.7 xhigh. From their docs: max: Max effort can deliver performance gains in some use cases, but may show diminishing returns from increased token usage. This setting…
I am on xhigh.