A question I've been asking alot lately (really since the release of GPT-5.3) is "do I really need the more powerful model"? I think a big issue with the industry right now is it's constantly chasing higher performing models and that comes at the cost of everything else. What I would love to see in the next few years is all these frontier AI labs go from just trying to create the most powerful model at any cost to ac…
Measuring Claude 4.7's tokenizer costs
401–410 of 540 posts
Re: Measuring Claude 4.7's tokenizer costs
#402I did some work yesterday with Opus and found it amazing. Today we are almost on non-speaking terms. I'm asking it to do some simple stuff and he's making incredible stupid mistakes: This is the third time that I have to ask you to remove the issue that was there for more than 20 hours. What is going on here? and at the same time the compacting is firing like crazy. (What adds ~4 minute delays every 1 - 15 minutes) |…
> he’s making .. mistakes Claude and other LLMs do not have a gender; they are not a “he”. Your LLM is a pile of weights, prompts, and a harness; anthropomorphising like this is getting in the way. You’re experiencing what happens when you sample repeatedly from a distribution. Given enough samples the probability of an eventual bad session is 100%. Just clear the context, roll back, and go again. This is part of the…
Re: Measuring Claude 4.7's tokenizer costs
#403Re: Measuring Claude 4.7's tokenizer costs
#404Earlier quoted context omitted.
I meant reference Toby Ord's work here. I think his framing of the performance/cost frontier hasn't gotten enough attention https://www.tobyord.com/writing/hourly-costs-for-ai-agents
That post doesn't address the human factor of cost, and I don't mean that in a good way. Even if AI costs more than a human, it's tireless, doesn't need holidays, is never going to have to go to HR for sexual harassment issues, won't show up hungover or need an advance to pay for a dying relative's surgery. It can be turned on and off with the flip of a switch. Hire 30 today, fire 25 of them next week. Spin another 5…
This is an architecture that people are increasing begging to give network connectivity that can't differentiate its system prompt from user input
Re: Measuring Claude 4.7's tokenizer costs
#405Earlier quoted context omitted.
So they nerfed 4.6 to make way for 4.7? Progress. /s
> they nerfed 4.6 to make way for 4.7? > Progress. /s pretty much, lmao. my theory is 4.6 started thinking less to save compute for 4.7 release. but who knows what's going on at anthropic
Re: Measuring Claude 4.7's tokenizer costs
#406Re: Measuring Claude 4.7's tokenizer costs
#407Earlier quoted context omitted.
> We need more voices like this to cut through the bullshit. Open models are not bullshit, they work fine for many cases and newer techniques like SSD offload make even 500B+ models accessible for simple uses (NOT real-time agentic coding!) on very limited hardware. Of course if you want the full-featured experience it's going to cost a lot.
I fell for this stuff, went into the open+local model rabbit hole, and am finally out of it. What a waste of time and money! People that love open models dramatically overstate how good the benchmaxxed open models are. They are nowhere near Opus.
I love my little hobby aquarium though... It's pretty impressive when Qwen Coder Next and Qwen 3.5 122B can accomplish (in terms of general agentic use and basic coding tasks), considering that the models are freely-available. (Also heard good things about Qwen 3.5 27B, but haven't used it much... yes I am a Qwen fanboi.)
Re: Measuring Claude 4.7's tokenizer costs
#408Re: Measuring Claude 4.7's tokenizer costs
#409Earlier quoted context omitted.
The cost to hire a human is highly predictable. The cost of AI isn't. I, as a human, need food and shelter, which puts a ceiling to my bargaining power. I can't withdraw my labour indefinitely. The power dynamics are also vastly against me. I represent a fraction of my employer's labour, but my employer represents 100% of my income. That dynamic is totally inverted with AI. You are a rounding error on their revenue s…
By continuously testing competitors and local LLMs? The reason for rising prices is that they (Anthropic) probably realized that they have reached a ceiling of what LLMs are capable of, and while it's a lot, it is still not a big moat and it's definitely not intelligence.
Re: Measuring Claude 4.7's tokenizer costs
#410I would rather steer quickly, get ideas because I'm moving quickly, do course correction quickly - basically I'm not happy blocking my chain of thought/concentration and fall prey to distractions due to Claude's slowness and compaction cycles. Sometimes I don't even notice that Codex has compacted.
For architectural discussions, sure I'll pick Claude. I'm mentally prepared for that. But once we are in the thick of things, speed matters. I would they rather focus on improving Sonnet's speed.