Live data from Hacker News

Measuring Claude 4.7's tokenizer costs

claudecodecamp.com

301–310 of 540 posts

Re: Measuring Claude 4.7's tokenizer costs

#301
post #29

IMHO there is a point where incremental model quality will hit diminishing returns. It is like comparing an 8K display to a 16K display because at normal viewing distance, the difference is imperceptible, but 16K comes at significant premium. The same applies to intelligence. Sure, some users might register a meaningful bump, but if 99% can't tell the difference in their day-to-day work, does it matter? A 20-30% cost…

I believe that's why 90% of the focus in these firms is on coding. There is a natural difficulty ramp-up that doesn't end anytime soon: you could imagine LLMs creating a line of code, a function, a file, a library, a codebase. The problem gets harder and harder and is still economically relevant very high into the difficulty ladder. Unlike basic natural language queries which saturate difficulty early. This is also w…

> the dimensionality of LLM output that is economically relevant keeps growing linearly for coding

Doubt. Yes. there was at one point it suddenly became useful to write code in a general sense. I have seen almost no improvement in department of architecting, operations and gaslighting. In fact gaslighting has gotten worse. Entire output based on wrong assumption that it hid, almost intentionally. And I had to create very dedicated, non-agentic tools to combat this.

And all of this with latest Opus line.

Re: Measuring Claude 4.7's tokenizer costs

#302
post #291

Earlier quoted context omitted.

They won't. These are not "issues", it's them trying to push the models to burn less compute. It will only get worse.

> it's them trying to push the models to burn less compute I'm curious, how does using more tokens save compute?

I think that the idea is each action uses more tokens, which means that users hit their limit sooner, and are consequently unable to burn more compute.

Re: Measuring Claude 4.7's tokenizer costs

#303

Earlier quoted context omitted.

I am having a shit experience lately. Opus 4.7, max effort. > You're right, that was a shit explanation. Let me go look at what V1 MTBL actually is before I try again. > Got it — I read the V1 code this time instead of guessing. Turns out my first take was wrong in an important way. Let me redo this in English. :facepalm:

The docs suggest not using max effort in most cases to avoid overthinking :shrug:

They've jumped the shark. I truly can't comprehend why all of these changes were necessary. They had a literal money printing machine that actually got real shit done, really well. Now it's a gamble every time and I am pulling back hard from Anthropic ecosystem.

Re: Measuring Claude 4.7's tokenizer costs

#304
post #291

Earlier quoted context omitted.

They won't. These are not "issues", it's them trying to push the models to burn less compute. It will only get worse.

> it's them trying to push the models to burn less compute I'm curious, how does using more tokens save compute?

It could be the adaptive reasoning

Re: Measuring Claude 4.7's tokenizer costs

#305

LLMs exist on a logaritmhic performance/cost frontier. It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost. To me, it is hard to reject this hypothesis today. The fact that Anthropic is rapidly trying to increase price may betray the fact that their recent lead is a…

> It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost. I think we're reaching the point where more developers need to start right-sizing the model and effort level to the task. It was easy to get comfortable with using the best model at the highest setting for every…

The problem is half the time you don't know you need the better model until the lesser model has made a massive mess. Then you have to do it again on the good model, wasting money. The "auto" modes don't seem to do a good job at picking a model IME.

Re: Measuring Claude 4.7's tokenizer costs

#306
post #205

Earlier quoted context omitted.

I would not be surprised at all, a $1,000/mo tool that makes your $20,000/mo engineer a lot more productive is an easy sell. I’m guessing we’re gonna have a world like working on cars - most people won’t have expensive tools (ex a full hydraulic lift) for personal stuff, they are gonna have to make do with lesser tools.

noway. i bought a $3k AMD395+ under the Sam Altman price hike and its got a local model that readily accomplishes medial tasks. theres a ceiling to these price hikes because open weights will keep popping up as competitors tey to advertise their wares. sure, we POV different capabilities but theres definitely not that much cash in propfietary models for their indererminance

[dead]

Re: Measuring Claude 4.7's tokenizer costs

#307
post #142

Earlier quoted context omitted.

Yep, and I just made a recommendation that was essentially "never enable Opus 4.7" to my org as a direct result. We have Opus 4.6 (3x) and Opus 4.5 (3x) enabled currently. They are worth it for planning . At 7.5x for 4.7, heck no. It isn't even clear it is an upgrade over Opus 4.6.

in copilot I find it hard to justify using opus at even 3x vs just using GPT 5.4 high at 1x

I went from plan with opus, implement with claude, to simply plan and implement with GPT 5.4

It's a very good model for a very good price

Re: Measuring Claude 4.7's tokenizer costs

#308

The title is a misdirection. The token counts may be higher, but the cost-per-task may not be for a given intelligence level. Need to wait to see Artificial Analysis' Intelligence Index run for this, or some other independent per-task cost analysis. The final calculation assumes that Opus 4.7 uses the exact same trajectory + reasoning output as Opus 4.6. I have not verified, but I assume it not to be the case, given…

I ran an internal (oil and gas focused) benchmark yesterday and found Opus 4.7 was 50% cheaper than Opus 4.6, driven by significantly fewer output tokens for reasoning. It also scored 80% (vs. 60%).

Re: Measuring Claude 4.7's tokenizer costs

#309

Earlier quoted context omitted.

> it's them trying to push the models to burn less compute I'm curious, how does using more tokens save compute?

I think that the idea is each action uses more tokens, which means that users hit their limit sooner, and are consequently unable to burn more compute.

What?

Re: Measuring Claude 4.7's tokenizer costs

#310
post #42

This is the backdoor way of raising prices... just inflate the token pricing. It's like ice cream companies shrinking the box instead of raising the price

No, you're forgetting the never ending world shattering models being released every couple of months. Each one with 2X token costs of course, for a vague performance gain and that will deprecate the previous ones.

It’s nice to see comments like this. It makes me feel less crazy. Something very weird is going on behind the scenes at Anthropic.
Post reply on HN