Live data from Hacker News

Measuring Claude 4.7's tokenizer costs

claudecodecamp.com

341–350 of 540 posts

Re: Measuring Claude 4.7's tokenizer costs

#341

A question I've been asking alot lately (really since the release of GPT-5.3) is "do I really need the more powerful model"? I think a big issue with the industry right now is it's constantly chasing higher performing models and that comes at the cost of everything else. What I would love to see in the next few years is all these frontier AI labs go from just trying to create the most powerful model at any cost to ac…

So you're happy with an untrustworthy lazy moron prone to stupid mistakes and guesswork?

Surely you can see the first lab that solves this gains a massive advantage?

Re: Measuring Claude 4.7's tokenizer costs

#342
post #289

Earlier quoted context omitted.

This seems like the experience I've had with every model I've tried over the last several years. It seems like an inherent limitation of the technology, despite the hyperbolic claims of those financially invested in all of this paying off.

Opus 4.6 pre-nerf was incredible, almost magical. It changed my understanding of how good models could be. But that's the only model that ever made me feel that way.

That was better, but still not to the point that I just let it go on my repo.

Re: Measuring Claude 4.7's tokenizer costs

#343
post #266

I did some work yesterday with Opus and found it amazing. Today we are almost on non-speaking terms. I'm asking it to do some simple stuff and he's making incredible stupid mistakes: This is the third time that I have to ask you to remove the issue that was there for more than 20 hours. What is going on here? and at the same time the compacting is firing like crazy. (What adds ~4 minute delays every 1 - 15 minutes) |…

> he’s making .. mistakes Claude and other LLMs do not have a gender; they are not a “he”. Your LLM is a pile of weights, prompts, and a harness; anthropomorphising like this is getting in the way. You’re experiencing what happens when you sample repeatedly from a distribution. Given enough samples the probability of an eventual bad session is 100%. Just clear the context, roll back, and go again. This is part of the…

Why be so upset at someone using pronouns with a LLM?

Re: Measuring Claude 4.7's tokenizer costs

#344
post #136

Earlier quoted context omitted.

I want to give give you realistic expectations: Unless you spend well over $10K on hardware, you will be disappointed, and will spend a lot of time getting there. For sophisticated coding tasks, at least. (For simple agentic work, you can get workable results with a 3090 or two, or even a couple 3060 12GBs for half the price. But they're pretty dumb, and it's a tease. Hobby territory, lots of dicking around.) Do your…

We need more voices like this to cut through the bullshit. It's fine that people want to tinker with local models, but there has been this narrative for too long that you can just buy more ram and run some small to medium sized model and be productive that way. You just can't, a 35b will never perform at the level of the same gen 500b+ model. It just won't and you are basically working with GPT-4 (the very first one…

> We need more voices like this to cut through the bullshit.

Just because you can't figure out how to use the open models effectively doesn't mean they're bullshit. It just takes more skill and experience to use them :)

Re: Measuring Claude 4.7's tokenizer costs

#345

LLMs exist on a logaritmhic performance/cost frontier. It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost. To me, it is hard to reject this hypothesis today. The fact that Anthropic is rapidly trying to increase price may betray the fact that their recent lead is a…

Once they implement their models directly in silicon, the cost will come down and the speed will go up. See Taalas.

Re: Measuring Claude 4.7's tokenizer costs

#346

Earlier quoted context omitted.

I believe that's why 90% of the focus in these firms is on coding. There is a natural difficulty ramp-up that doesn't end anytime soon: you could imagine LLMs creating a line of code, a function, a file, a library, a codebase. The problem gets harder and harder and is still economically relevant very high into the difficulty ladder. Unlike basic natural language queries which saturate difficulty early. This is also w…

> the dimensionality of LLM output that is economically relevant keeps growing linearly for coding Doubt. Yes. there was at one point it suddenly became useful to write code in a general sense. I have seen almost no improvement in department of architecting, operations and gaslighting. In fact gaslighting has gotten worse. Entire output based on wrong assumption that it hid, almost intentionally. And I had to create…

Also doubt. But most likely because of organizational inertia. After a while, you’re mostly focused on small problems and big features are rare. You solution is quasi done. But now each new change is harder because you don’t want to broke assumptions that have become hard requirements.

Re: Measuring Claude 4.7's tokenizer costs

#347
post #3

Earlier quoted context omitted.

haven't people been complaining lately about 4.6 getting worse?

People complain about a lot of things. Claude has been fine: https://marginlab.ai/trackers/claude-code-historical-perform...

Your link shows there have been huge drops.

How is it fine?

Re: Measuring Claude 4.7's tokenizer costs

#348

Earlier quoted context omitted.

Its enshittificating real fast. They'll just keep releasing model after model, more expensive than the last, marginal gains, but touted as "the next thing". Evangelists will say that they're afraid, it's the future, in 6 months it's all over. Anthropic will keep astroturfing on Reddit. CEOs will make even more outlandish claims. You raised a good point, what's a good metric for LLM performance? There's surely all the…

This is most likely trajectory I fear. It reminds me a lot of Oracle, where they rebrand and reskin products just to change pricing/marketing without adding anything.

Win 10, win 11, all the recent macOS,… could have been released as features and not new products

Re: Measuring Claude 4.7's tokenizer costs

#349

Earlier quoted context omitted.

> It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost. I think we're reaching the point where more developers need to start right-sizing the model and effort level to the task. It was easy to get comfortable with using the best model at the highest setting for every…

The problem is half the time you don't know you need the better model until the lesser model has made a massive mess. Then you have to do it again on the good model, wasting money. The "auto" modes don't seem to do a good job at picking a model IME.

[deleted]

Re: Measuring Claude 4.7's tokenizer costs

#350
post #42

This is the backdoor way of raising prices... just inflate the token pricing. It's like ice cream companies shrinking the box instead of raising the price

No, you're forgetting the never ending world shattering models being released every couple of months. Each one with 2X token costs of course, for a vague performance gain and that will deprecate the previous ones.

https://platform.claude.com/docs/en/about-claude/model-depre...

Retirement date for Opus 4.6 is marked as "Not sooner than February 5, 2027"

Post reply on HN