Earlier quoted context omitted.
I have a project where we've had Opus, Sonnet, Deepseek, Kimi, Qwen create and execute an aggregate total of about 350 plans so far, and the quality difference as measured in plans where the agent failed to complete the tasks on the first run is high enough that it comes out several times higher than Anthropics subscription prices, but probably cheaper than the API prices once we have improved the harness further - a…
In 12 months, opus will be better than now and you still won't use it lol
My default model has now dropped to Sonnet, because Sonnet can now do most of my tasks, and we already use Kimi, Deepseek, and Qwen.
They're just not cost-effective enough to be my main driver yet. They are however cheap enough that for things where the Claude TOS does not let me use my subscription, they still add substantial value. Just not nearly as much as I'd like.
The bulk of my tasks won't get harder as time passes, and so will move down the value chain as the cheaper models get better.
For the small proportion of my tasks that benefits from a smarter model, I will use the smartest model I can afford.