I thought this was common knowledge. DeepSeek’s Wikipedia entry says that they trained all their models on Nvidia chips procured before the U.S. embargo to China on them. It wouldn’t surprise me if they continued acquiring them through, well, less than legal means. I also read somewhere (not Wikipedia) that they trained on ChatGPT, Claude, and Gemini queries, basically feeding in the output of competitor’s LLMs as tr…
Not sure why you would expect this, all the models started doing this as its much more cost effective to get data for post training don't you remember the first grok release where many times it started replies "as a model trained by openai..."