Regarding LLMs we're in a race to the bottom. Chinese models perform similarly with much higher efficiency; refer to kimi-k2 and plenty of others. ClopenAI is extremely overvalued, and AGI is not around the corner because among 20T+ tokens trained on it still generates 0 novel output. Try asking for ASP.NET Core .MapOpenAPI() instead of the pre .net9 swashbuckle version. You get nothing. It's not in the training data…
They perform similarly on benchmarks, which can be fudged to arbitrarily high numbers by just including the Q&A into the training data at a certain frequency or post-training on it. I have not been impressed with any of the DeepSeek models in real-world use.
Anecdata: our product is using a number of these models in production.