Earlier quoted context omitted.
That's what I mean, those CoT are never ending currently until you run out of context.
I'm not sure if you're talking conversationally and I'm taking it as a technical query, or you're saying CoT never terminate for you and asking for input, or asking what the paper implies about CoT, or relaying that you understand the papers claim that this method net reduces CoT length.
Currently it seems that this shift of cost to inference-time is a necessary tradeoff that we have to live with (for now at least).
It's more of a problem for many people running those models locally because they have constrained hardware that can't handle those long contexts.
It seems that this paper shows not only that their method is cheaper in terms of fine tuning but also significantly reduces inference time cost for CoTs.
If what they say gets confirmed it looks to me like quite significant contribution?