The incredible arrogance and hybris of the American initiated tech war - it is just a beautiful thing to see it slowly fall apart. The US-China contest aside - it is in the application layer llms will show their value. There the field, with llm commoditization and no clear monopolies, is wide open. There was a point in time where it looked like llms would the domain of a single well guarded monopoly - that would have…
As much I apprecite the sentiment, I think it is too early to declare that the well guareded monopoly is over. Yes, these models have answers, but don't expect all the large enterprises to switch to these models. The other aspect is scaling to serve these models will need a lot of time even if Huawei succeeds. Not all the Governments trust China and there will be a lot of resistance to work with these models eventual…
DeepSeek v4
641–650 of 1001 posts
Re: DeepSeek v4
#642I just did some quick testing on my own benchmark that tests LLMs as customer support chatbots, and found out that deepseek-v4-flash (scored 90.2%) was better than qwen3.5-27b (89%) and qwen3.5-35b-a3b (89.1%) and roughly equal to gemini-3-flash-preview (90.5%), but deepseek-v4-flash had the lowest cost of all of them by far. Half the cost of gemini-3-flash and an order of magnitude less cost than the qwen models. Ha…
It's five times bigger in both total and active parameters!
Re: DeepSeek v4
#643Re: DeepSeek v4
#644There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…
https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 https://huggingface.co/deepseek-ai/DeepSeek-Prover-V2-671B
Re: DeepSeek v4
#645Wow, never seen a post with so many comments posted overnight like this.
Re: DeepSeek v4
#646There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…
Very interesting. I wonder how much of this is due to the context length. I am unclear on the implementation strategy, you ran this problem as a 1-shot using chat mode, or using each on an agent harness?
Re: DeepSeek v4
#647Earlier quoted context omitted.
Hmm, the Flash performs significantly better than Pro in the benchmark? That's very strange; could rate limiting cause that?
Yes, Flash doesn't seem to have the same rate limits as Pro. I expect once the API issues are fixed, for v4-pro to be around the same level as GLM-5.
(I am confused by the results your website is presenting)
Re: DeepSeek v4
#648Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…
The funniest thing is how Americans have been fooled with this stuff. This version of AI is mostly taking a public paper from 2017, investing in GPUs, and feeding it as much data as possible. So with a few computer scientists, no respect for intellectual property, and tons of money to burn, you have all the ingredients to create this technology. Sam Altman and friends did it, as did the Chinese. The difference is tha…
Re: DeepSeek v4
#649This is shockingly cheap for a near frontier model. This is insane. For context, for an agent we're working on, we're using 5-mini, which is $2/1m tokens. This is $0.30/1m tokens. And it's Opus 4.6 level - this can't be real. I am uncomfortable about sending user data which may contain PII to their servers in China so I won't be using this as appealing as it sounds. I need this to come to a US-hosted environment at a…
> I am uncomfortable about sending user data which may contain PII to their servers in China As a European I feel deeply uncomfortable about sending data to US companies where I know for sure that the government has access to it. I also feel uncomfortable sending it to China. If you'd asked me ten years ago which one made me more uncomfortable. China. But now I'm not so sure, in fact I'm starting to lean towards the…
Re: DeepSeek v4
#650Earlier quoted context omitted.
Jensen Huang said this in his recent interview - that China has the best/most engineers, it has the chip making ability, it's a good thing they wanna build on a Nvidia stack - but if you push them they will build on an all Chinese stack - but the interviewer was being a numb head who kept parroting the propaganda of Western tech supremacy
Referring to the Dwarkesh interview clearly. Jensen came across as incredibly defensive and intentionally close-minded, shows that even billionaires suffer from "a man can't understand something if his paycheck depends on him not understanding it." Your assertion is silly: did Tesla selling electric cars into China stop them from delivering their own industry? They were going to develop their domestic industry regard…