Is there a general solution to this problem? I assume you can only start buffering tokens once you see a construct, for which there are continuations, that once completed, would lead to the previous text being rendered differently. Of course you don't want to keep buffering for too long, since this would defeat the purpose of streaming. And you never know if the potential construct will actually be generated. Also, t…
Almost every model has a slight but meaningfully different opinion on what markdown is and how creative they can be with it.
Doing it well is a non-trivial problem.