The title is a misdirection. The token counts may be higher, but the cost-per-task may not be for a given intelligence level. Need to wait to see Artificial Analysis' Intelligence Index run for this, or some other independent per-task cost analysis. The final calculation assumes that Opus 4.7 uses the exact same trajectory + reasoning output as Opus 4.6. I have not verified, but I assume it not to be the case, given…
Measuring Claude 4.7's tokenizer costs
161–170 of 540 posts
Re: Measuring Claude 4.7's tokenizer costs
#162But it looks like it's just creeping up. Probably because we're paying for construction, not just inference right now.
Re: Measuring Claude 4.7's tokenizer costs
#163Re: Measuring Claude 4.7's tokenizer costs
#164A question I've been asking alot lately (really since the release of GPT-5.3) is "do I really need the more powerful model"? I think a big issue with the industry right now is it's constantly chasing higher performing models and that comes at the cost of everything else. What I would love to see in the next few years is all these frontier AI labs go from just trying to create the most powerful model at any cost to ac…
Re: Measuring Claude 4.7's tokenizer costs
#165Re: Measuring Claude 4.7's tokenizer costs
#166Why release this?
Re: Measuring Claude 4.7's tokenizer costs
#167Earlier quoted context omitted.
The problem is that people equate money to power and power to evil. So no matter what, if you do something lots of people like (and hence compensate you for), you will be evil. It's a very interesting quirk of human intuition.
A reasonable conclusion, considering that money and power seem to have their own gravity, so people with more of both end up getting even more of both, and vice versa. Can't blame someone who comes to such a conclusion about money and power.
Re: Measuring Claude 4.7's tokenizer costs
#168Re: Measuring Claude 4.7's tokenizer costs
#169IMHO there is a point where incremental model quality will hit diminishing returns. It is like comparing an 8K display to a 16K display because at normal viewing distance, the difference is imperceptible, but 16K comes at significant premium. The same applies to intelligence. Sure, some users might register a meaningful bump, but if 99% can't tell the difference in their day-to-day work, does it matter? A 20-30% cost…
yeah thats is my biggest issue - im okay with paying 20-30% more but what is the ROI? i dont see an equivalent improvement in performance. Anthropic hasnt published any data around what these improvements are - just some vague “better instruction following"
You raised a good point, what's a good metric for LLM performance? There's surely all the benchmarks out there, but aren't they one and done? Usually at release? What keeps checking the performance of those models. At this point it's just by feel. People say models have been dumbed down, and that's it.
I think the actual future is open source models. Problem is, they don't have the huge marketing budget Anthropic or OpenAI does.