Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

321–330 of 1001 posts

Re: DeepSeek v4

#322
The incredible arrogance and hybris of the American initiated tech war - it is just a beautiful thing to see it slowly fall apart.

The US-China contest aside - it is in the application layer llms will show their value. There the field, with llm commoditization and no clear monopolies, is wide open.

There was a point in time where it looked like llms would the domain of a single well guarded monopoly - that would have been a very dark world. Luckily we are not there now and there is plenty of grounds for optimism.

Re: DeepSeek v4

#323
post #244

Earlier quoted context omitted.

It's because they're optimizing for a different problem. Western Models are optimizing to be used as an interchangeable product. Chinese models are being optimizing to be built upon.

[flagged]

They are all racing to AGI. They aren't designing them to be interchangeable they just happen to be.

Re: DeepSeek v4

#324

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

Curious to know what kind of problems you are talking about here

I don't want to give away too much due to anonymity reasons, but the problems are generally in the following areas (in order from hardest to easiest):

- One problem on using quantum mechanics and C*-algebra techniques for non-Markovian stochastic processes. The interchange between the physics and probability languages often trips the models up, so pretty much everything tends to fail here.

- Three problems in random matrix theory and free probability; these require strong combinatorial skills and a good understanding of novel definitions, requiring multiple papers for context.

- One problem in saddle-point approximation; I've just recently put together a manuscript for this one with a masters student, so it isn't trivial either, but does not require as much insight.

- One problem pertaining to bounds on integral probability metrics for time-series modelling.

Re: DeepSeek v4

#325
post #77

For comparison on openrouter DeepSeek v4 Flash is slightly cheaper than Gemma 4 31b, more expensive than Gemma 4 26b, but it does support prompt caching, which means for some applications it will be the cheapest. Excited to see how it compares with Gemma 4.

I wonder why there aren't more open weights model with support for prompt caching on OpenRouter.

It is tricky to build good infrastructure for prompt caching.

Re: DeepSeek v4

#326

Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…

I sometimes wonder if there are any security risks with using Chinese LLMs. Is there?

Theoretically yes. It is entirely possible to poison the training data for a supply chain attack against vibe coders. The trick would be to make it extremely specific for a high value target so it is not picked up by a wide range of people. You could also target a specific open source project that is used by another widely used product.

However there is so many factors involved beyond your control that it would not be a viable option compared to other possible security attacks.

Re: DeepSeek v4

#327
post #311

Earlier quoted context omitted.

As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed

"not perfect" is a _very_ big simplification of what China is though

Isn't that the same to every major superpower?

Re: DeepSeek v4

#328
post #281

So, this is the version that's able to serve inference from Huawei chips, although it was still trained on nVidia. So unless I'm very much mistaken this is the biggest and best model yet served on (sort of) readily-available chinese-native tech. Performance and stability will be interesting to see; openrouter currently saying about 1.12s and 30tps, which isn't wonderful but it's day one after all. For reference, the…

Great! Can't wait to buy decent GPU for interference for <1k$

Re: DeepSeek v4

#329
post #311

Earlier quoted context omitted.

As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed

"not perfect" is a _very_ big simplification of what China is though

You can say the same about the US

Re: DeepSeek v4

#330

Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…

As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed

I don’t know if we’re ahead of the curve but that tired feeling has started turning into hate here in the EU. I guess being threatened with invasion does that to you.

The next decade is going to look very different with America Alone.

Post reply on HN