Earlier quoted context omitted.
nobody has shown CoT scaling like this except deepmind, it is very obviously a result of their alignment pipeline not just prompting.
Scaling like what? Are there any comparisons with and without CoT, or with other models with their CoT? As far as I'm aware, their CoT part is secret. I'm sure the finetuning does some lifting, but I'm also sure the difference in a fair comparison won't be remotely as significant as it's being hyped currently. This is still clearly CoT, with all its limitations and caveats as expected. That's an improvement, sure, bu…
Saying it's just CoT is kind of meaningless. Even just looking at the examples on Open AI's blog and you quickly see no other model today can generate or utilize CoT to anywhere near that quality through prompting or naive fine-tuning.