LLM pricing makes no sense. Why pay the same rate for tokens that are in/out of the context window when that's 99% of performance.

Now the burden is on the user to somehow guess or reverse-engineer the internals of your model