Earlier quoted context omitted.
I think it's fairer to compare it to the original GPT-4 which might the equivalent in term of "size" (though we don't have actual numbers for either). GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens So 4.5 is 2.5x more expensive. I think they announced this as their last non-reasoning model, so it was maybe with the goal of stretching pre-training as far as they could, just to see what new capabilities wo…
Why would that be fairer? We can assume they did incorporate all learnings and optimizations they made post gpt-4 launch, no?
It'd be weird to release a distilled version without ever releasing the base undistilled version.