Live data from Hacker News

Measuring Claude 4.7's tokenizer costs

claudecodecamp.com

231–240 of 540 posts

Re: Measuring Claude 4.7's tokenizer costs

#232
post #21

Yeah. I just did a day with 4.7 and I won't be going back for a while. It is just too expensive. On top of the tokenization the thinking seems like it is eating a lot more too.

yeah i am still not clear why there are 5 effort modes now on top of more expensive tokenization

choice is often a great dark-pattern (lack of choice is too but...). Choices generally grow cost to discover optimality in an np way. This means if the entity giving choice has more ability to compute the value prop than the entity deciding the choice you can easily create an exploitive system. Just create a bunch of choices, some actually do save money with enough thought but most don't, and you will gain:

People that think they got what they wanted, the feature is there!, so they can't complain but...

People that end up essentially randomly picking so the average value of the choices made by customers is suboptimal.

Re: Measuring Claude 4.7's tokenizer costs

#233

A question I've been asking alot lately (really since the release of GPT-5.3) is "do I really need the more powerful model"? I think a big issue with the industry right now is it's constantly chasing higher performing models and that comes at the cost of everything else. What I would love to see in the next few years is all these frontier AI labs go from just trying to create the most powerful model at any cost to ac…

Yes! I'd be totally happy with today's sonnet 4.6 if I could run it locally.

If you can forgive the obviously-AI-generated writing, [CPUs Aren't Dead](https://seqpu.com/CPUsArentDead) makes an interesting point on AI progress: Google's latest, smallest Gemma model (Gemma 4 E2B), which can run on a cell phone, outperforms GPT-3.5-turbo. Granted, this factoid is based on `MT-Bench` performance, a benchmark from 2023 which I assume to be both fully saturated and leaked into the training data for modern LLMs. However, cross-referencing [Artificial Analysis' Intelligence Index](https://artificialanalysis.ai/models?models=gemma-4-e2b-non-...) suggests that indeed the latest 2B open-weights models are capable of matching or beating 175B models from 3-4 years ago. Perhaps more impressive, [Gemma 4 E4B matches or beats GPT-4o](https://artificialanalysis.ai/models?models=gemma-4-e4b%2Cge...) on many benchmarks.

If this trend continues, perhaps we'll have the capabilities of today's best models available to reasonably run on our laptops!

Re: Measuring Claude 4.7's tokenizer costs

#234
I've been using 4.6 models since each of them launched. Same for 4.5.

4.6 performers worse or the same in most of the tasks I have. If there is a parameter that made me use 4.6 more frequently is because 4.5 get dumber and not because 4.6 seemed smarter.

Re: Measuring Claude 4.7's tokenizer costs

#235
In Kolkata, sweet sellers was struggling with cost management after covid due to increased prices of raw materials. But they couldn't increase the price any further without losing customers. So they reduced the size of sweets instead, and market slowly reduced expectations. And this is the new normal now.

Human psychology is surprisingly similar, and same pattern comes across domains.

Re: Measuring Claude 4.7's tokenizer costs

#237

Claude's tokenizers have actually been getting less efficient over the years (I think we're at the third iteration at the least since Sonnet 3.5). And if you prompt the LLM in a language other than English, or if your users prompt it or generate content in other languages, the costs go higher even more. And I mean hundreds of percent more for languages with complex scripts like Tamil or Japanese. If you're interested…

I would encourage you to post a link here, and also to submit to HN if you haven't already. :)

Will do! Thanks for the encouragement

Re: Measuring Claude 4.7's tokenizer costs

#238

Earlier quoted context omitted.

Does anyone here use 8k display for work? Does it make sense over 4k? I was always wondering where that breaking point for cost/peformance is for displays. I use 4K 27” and it’s noticeably much better for text than 1440p@27 but no idea if the next/ and final stop is 6k or 8k?

Even 4k turns out to be overkill if you're looking at the whole screen and a pixel-perfect display. By human visual acuity, 1440p ought to be enough, and even that's taking a safety margin over 1080p to account for the crispness of typical text.

1440p is enough if you haven't experienced anything else. Even the jump from 4k to 5-6k is quite noticeable on a 27" monitor.

I switched to the Studio Display XDR and it is noticeably better than my 4k displays and my 1440p displays feel positively ancient and near unusable for text.

Re: Measuring Claude 4.7's tokenizer costs

#239

The "multiplier" on Github Copilot went from 3 to 7.5. Nice to see that it is actually only 20-30% and Microsoft wanting to lose money slightly slower. https://docs.github.com/fr/copilot/reference/ai-models/suppo...

Yep, and I just made a recommendation that was essentially "never enable Opus 4.7" to my org as a direct result. We have Opus 4.6 (3x) and Opus 4.5 (3x) enabled currently. They are worth it for planning . At 7.5x for 4.7, heck no. It isn't even clear it is an upgrade over Opus 4.6.

I don't know how you guys are not seeing 4.7 as an upgrade, it just does so much more, so much better. I guess lower complexity tasks are saturated though.

Re: Measuring Claude 4.7's tokenizer costs

#240

In Kolkata, sweet sellers was struggling with cost management after covid due to increased prices of raw materials. But they couldn't increase the price any further without losing customers. So they reduced the size of sweets instead, and market slowly reduced expectations. And this is the new normal now. Human psychology is surprisingly similar, and same pattern comes across domains.

It's not just in Kolkata, worldwide packs of biscuits etc remained the same size but less inside.

I didn't buy Springles chips in years, even the box now is nothing like it was. Thinner. Shorter. I imagine how far from the top the slices stack up.

Post reply on HN