Earlier quoted context omitted.
> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the training data. ...and because that training data is missing, they can't be replicated. Which means that you cannot assert that the Chinese are being open in their LLM development, because there's no way to verify that the techniques they describe are actually the…
You can replicate the architectural innovations, and try them for yourself with your own dataset. It seems some of them are certainly being used by western companies, such as DeepSeek Sparse Attention, now supported by NVIDIA cuDNN. Ditto for training algorithms and procedures such as Slime or DeepSeek's details instructions on how to build a reasoning model. This is the exact value of openly shared details - others…
That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data.
> and the reason the Chinese are not sharing data are no more nefarious than why the American companies are not sharing
This is moving the goalposts. Your claim was that "The Chinese have actually been very open about training", which is false, as discussed. Nobody ever claimed that the American labs were open.