Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

641–650 of 1001 posts

Re: DeepSeek v4

#641
post #322

The incredible arrogance and hybris of the American initiated tech war - it is just a beautiful thing to see it slowly fall apart. The US-China contest aside - it is in the application layer llms will show their value. There the field, with llm commoditization and no clear monopolies, is wide open. There was a point in time where it looked like llms would the domain of a single well guarded monopoly - that would have…

As much I apprecite the sentiment, I think it is too early to declare that the well guareded monopoly is over. Yes, these models have answers, but don't expect all the large enterprises to switch to these models. The other aspect is scaling to serve these models will need a lot of time even if Huawei succeeds. Not all the Governments trust China and there will be a lot of resistance to work with these models eventual…

Which Monopoly? Are all large enterprises in USA? There are tons of them outside and they will run the open ones and cheapest ones to infer and those are Chinese. I run Chinese models at home and don't bother with cloud. If I could call the shots at work, we will switch 100% to Chinese models so everyone could have "unlimited" tokens.

Re: DeepSeek v4

#642

I just did some quick testing on my own benchmark that tests LLMs as customer support chatbots, and found out that deepseek-v4-flash (scored 90.2%) was better than qwen3.5-27b (89%) and qwen3.5-35b-a3b (89.1%) and roughly equal to gemini-3-flash-preview (90.5%), but deepseek-v4-flash had the lowest cost of all of them by far. Half the cost of gemini-3-flash and an order of magnitude less cost than the qwen models. Ha…

How can a medium-sized model like Deepseek-V4-Flash be cheaper than a much smaller models like Qwen3.5-35B-A3B.

It's five times bigger in both total and active parameters!

Re: DeepSeek v4

#644

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

They have had the best math models for about a year most folks just didn't know about it. You can't find inference on APIs, but I run these at home, this is also the advantage of open models.

https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 https://huggingface.co/deepseek-ai/DeepSeek-Prover-V2-671B

Re: DeepSeek v4

#646

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

Very interesting. I wonder how much of this is due to the context length. I am unclear on the implementation strategy, you ran this problem as a 1-shot using chat mode, or using each on an agent harness?

Has nothing to do with context length, they have experience training math models, they have a model that would take gold in IMO and a lean prover. Both have been out for almost a year.

Re: DeepSeek v4

#647
post #621
post #612

Earlier quoted context omitted.

Hmm, the Flash performs significantly better than Pro in the benchmark? That's very strange; could rate limiting cause that?

Yes, Flash doesn't seem to have the same rate limits as Pro. I expect once the API issues are fixed, for v4-pro to be around the same level as GLM-5.

Why would your test be including scores of failed responses/runs? That seems confusing.

(I am confused by the results your website is presenting)

Re: DeepSeek v4

#648

Open Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when…

The funniest thing is how Americans have been fooled with this stuff. This version of AI is mostly taking a public paper from 2017, investing in GPUs, and feeding it as much data as possible. So with a few computer scientists, no respect for intellectual property, and tons of money to burn, you have all the ingredients to create this technology. Sam Altman and friends did it, as did the Chinese. The difference is tha…

[deleted]

Re: DeepSeek v4

#649

This is shockingly cheap for a near frontier model. This is insane. For context, for an agent we're working on, we're using 5-mini, which is $2/1m tokens. This is $0.30/1m tokens. And it's Opus 4.6 level - this can't be real. I am uncomfortable about sending user data which may contain PII to their servers in China so I won't be using this as appealing as it sounds. I need this to come to a US-hosted environment at a…

> I am uncomfortable about sending user data which may contain PII to their servers in China As a European I feel deeply uncomfortable about sending data to US companies where I know for sure that the government has access to it. I also feel uncomfortable sending it to China. If you'd asked me ten years ago which one made me more uncomfortable. China. But now I'm not so sure, in fact I'm starting to lean towards the…

The chances of my bank account getting hacked due to the PLA backdoor in Deepseek is higher than the CIA backdoor in OpenAI.

Re: DeepSeek v4

#650
post #521

Earlier quoted context omitted.

Jensen Huang said this in his recent interview - that China has the best/most engineers, it has the chip making ability, it's a good thing they wanna build on a Nvidia stack - but if you push them they will build on an all Chinese stack - but the interviewer was being a numb head who kept parroting the propaganda of Western tech supremacy

Referring to the Dwarkesh interview clearly. Jensen came across as incredibly defensive and intentionally close-minded, shows that even billionaires suffer from "a man can't understand something if his paycheck depends on him not understanding it." Your assertion is silly: did Tesla selling electric cars into China stop them from delivering their own industry? They were going to develop their domestic industry regard…

I thought Jensen’s comparison to Huawei’s cell phone hardware infra (towers and networking) to be an interesting comparison- that shutting them out of a market was one of the causes of their current position in the market. It made them more dominant in the end.
Post reply on HN