Cool! I love seeing the very subtle emergent behaviours of different CoT approaches. I reckon people still don't fully appreciate the brittle artistic subtlety of trying to make something akin to logic emerge from these weird transformer machines. > Now, instead asking the model to respond in JSON we use XML tags to separate the start and end of the contemplation phase and the final answer I suspect the author wisely…
I always found the “structured output feature” thing odd. LLMs will follow any structure present in the prompt. So with a few few-shot examples, you have always been able to make them do anything - return JSON, Python function calls, or any syntax you come up with that works well for your task (ie streaming, separating thinking tokens from output, etc.).
All of that comes just with a little bit of overhead for the developer: Giving the examples, writing the parsing logic, and a strategy for the rare cases where the parsing may fail (retry, default value, prompt adaptation).
Once you embrace this approach, switching models is trivial. You can evaluate models with heavy RLHF like OpenAI’s against open models, base models, tiny models. Very easy to save a lot of money and achieve incredible speeds.