The real prices of frontier models
playcode.io
The real prices of frontier models
1–10 of 91 posts
Re: The real prices of frontier models
#2Chattiness remains an open issue for some of the SoTA open weights & (to a lesser extent) Claude.
Re: The real prices of frontier models
#3Re: The real prices of frontier models
#4Other traits where models differ that have an even greater impact on your total spend:
* How much context do they load in to solve a given task?
* How long do they spend thinking to get equivalent results?
* How many times do they stop and ask you for input, and are you there to respond to them before the cache runs out?
* Etc.
Incorporating the tokenizer just makes a very imprecise measurement of cost a little bit more precise, but in my own experience I have not found that the token cost is a significant driver of task cost whether or not you incorporate the tokenizer. Everything else about the model's behavior has a much larger impact.
Re: The real prices of frontier models
#5Re: The real prices of frontier models
#6Tokenzier aside, a report shared on reddit found that the GPT 5.6 (edit: 5.5) series are incredibly thrifty with CoTs, resulting in cheaper bills than GLM 5.2 (let alone Opus/Fable): https://www.reddit.com/r/ZaiGLM/s/rUoG5adkPh Chattiness remains an open issue for some of the SoTA open weights & (to a lesser extent) Claude.
Re: The real prices of frontier models
#7Re: The real prices of frontier models
#8- A ~2000-2002 legacy C++ game codebase at about ~90kloc: GPT 1.12M, Claude 2.2M
- A ~30kloc TypeScript codebase: GPT 260K, Claude 437K
In the end, GPT's current tokenizer is ~1.6x-2x better than Claude's current one, depending on your data. And you can check for free for both, for OpenAI just use the open-source libraries, for Anthropic - you have to use their count_tokens endpoint as they don't publish the tokenizer, but the endpoint is free (and allows requests over 1M tokens as well).