Earlier quoted context omitted.
But remember to not ask about Taiwan!
Quit a bit better then made to bomb little girl schools in Iran.
DeepSeek v4
321–330 of 1001 posts
Re: DeepSeek v4
#322The US-China contest aside - it is in the application layer llms will show their value. There the field, with llm commoditization and no clear monopolies, is wide open.
There was a point in time where it looked like llms would the domain of a single well guarded monopoly - that would have been a very dark world. Luckily we are not there now and there is plenty of grounds for optimism.
Re: DeepSeek v4
#323Earlier quoted context omitted.
It's because they're optimizing for a different problem. Western Models are optimizing to be used as an interchangeable product. Chinese models are being optimizing to be built upon.
[flagged]
Re: DeepSeek v4
#324There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…
Curious to know what kind of problems you are talking about here
- One problem on using quantum mechanics and C*-algebra techniques for non-Markovian stochastic processes. The interchange between the physics and probability languages often trips the models up, so pretty much everything tends to fail here.
- Three problems in random matrix theory and free probability; these require strong combinatorial skills and a good understanding of novel definitions, requiring multiple papers for context.
- One problem in saddle-point approximation; I've just recently put together a manuscript for this one with a masters student, so it isn't trivial either, but does not require as much insight.
- One problem pertaining to bounds on integral probability metrics for time-series modelling.
Re: DeepSeek v4
#325For comparison on openrouter DeepSeek v4 Flash is slightly cheaper than Gemma 4 31b, more expensive than Gemma 4 26b, but it does support prompt caching, which means for some applications it will be the cheapest. Excited to see how it compares with Gemma 4.
I wonder why there aren't more open weights model with support for prompt caching on OpenRouter.
Re: DeepSeek v4
#326Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…
I sometimes wonder if there are any security risks with using Chinese LLMs. Is there?
However there is so many factors involved beyond your control that it would not be a viable option compared to other possible security attacks.
Re: DeepSeek v4
#327Earlier quoted context omitted.
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
"not perfect" is a _very_ big simplification of what China is though
Re: DeepSeek v4
#328So, this is the version that's able to serve inference from Huawei chips, although it was still trained on nVidia. So unless I'm very much mistaken this is the biggest and best model yet served on (sort of) readily-available chinese-native tech. Performance and stability will be interesting to see; openrouter currently saying about 1.12s and 30tps, which isn't wonderful but it's day one after all. For reference, the…
Re: DeepSeek v4
#329Earlier quoted context omitted.
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
"not perfect" is a _very_ big simplification of what China is though
Re: DeepSeek v4
#330Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
The next decade is going to look very different with America Alone.