Live data from Hacker News

DeepSeek-V4-Flash Update

api-docs.deepseek.com

221–230 of 362 posts

Re: DeepSeek-V4-Flash Update

#221

Earlier quoted context omitted.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

Reuters was reporting that rumor. And then Xi made a public appearance at a conference in Shanghai where he said the opposite of that rumor.

No in that very speech Xi stated explicitly what amounted to: 'of course when we get a Mythos, it will be a state secret'.

In fact the overwhelming weight of AI use in China, the chatgpt so to say, is Bytedance's AI which is absolutely closed and uniquely opaque.

The press treatment of these matters was no good and they are slowly walking it back, e.g. NYT yesterday finally actually read the speech.

Re: DeepSeek-V4-Flash Update

#222

Earlier quoted context omitted.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold

Even Xi's speech was as paranoid as Dario Amodei about the possibility of Chinese AI achieving something at the level of say Mythos. If you seriously believe such a thing will be on Hugging Face I really don't know what to say.

Re: DeepSeek-V4-Flash Update

#223
post #204

Earlier quoted context omitted.

What makes you state this?

Decades of xenophobic propaganda

No, reading Xi's speech. They aren't going to give their Mythos - presumably a few months away - to the Sinaloa cartel or Uighur hackers or the US military or etc etc

Re: DeepSeek-V4-Flash Update

#224

Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc

It's been good for one-off cases in my limited experience. I would describe its behaviour as Sonnet-shaped, if that makes sense to you. Good answers but it often decides to reason a lot about simple things before getting to an output.

At the speed Flash has on most providers, it doesn't really turn into a latency concern.

Re: DeepSeek-V4-Flash Update

#225

Earlier quoted context omitted.

What makes you state this?

China bad

I'm a) quoting Xi's speech and b) projecting rapid China development to Mythos level. Either you think China will never get to 'Mythos level' because you think China bad; or you think they will release the weights, because you think China bad. Its incredible the weight of ideology over facts in this discourse

Re: DeepSeek-V4-Flash Update

#226

Is anyone using DSv4 for their agents that is not related to writing code? Curious about use cases specially for someone using gpt-5.4 mini for classification, categorization, etc

We use it at a moderate scale, self-hosted on B300 hardware. It's great :D

QA analysis of voice transcriptions. Napkin math: we operate at 2-5% of the cost of running on Equiv Frontier, though this changes near-weekly because pricing is so volatile.

It took us about a month to get the inference configured to achieve these numbers. But if you can get your hands on a pair of B300 GPUs and the context works, it's untouchable for price/performance.

(B200 would work, but you don't have the B300's memory, which lets you run it on 2xGPU instead of 4xGPU... with Dspark, it's like magic)

On a side note, for tasks that don't require the intelligence of DS v4 flash, we're using Nemotron-3-super with incredible success. I'm shocked we're not seeing more adoption of this model, given how easy it is to fine-tune and how blisteringly fast the nvfp4 version is. (A single B200 GPU can produce an insane amount of throughput with Nemotron 3 Super.)

Re: DeepSeek-V4-Flash Update

#227
post #220

Earlier quoted context omitted.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

We can only wait to see if this is true. Another possibility could be that it was an answer to USA's government considering a ban on chinese models. In this way, he fueled the discussion around the importance of open weight models.

Xi actually said that models of the strength to pose security issues will be permanent state secrets -- in the same speech that the first western takes translated as actually using the words 'open weights' which of course nowhere appeared. Now: when will "models of the strength to pose security issues" appear? Maybe my inference 'months' is wrong, but they seem to be advancing rapidly.

The administration blather about banning open weights is characteristically confused. Xi has already stated (what is obvious) that he will ban security-endangering weights and keep them a state secret.

Re: DeepSeek-V4-Flash Update

#228

Earlier quoted context omitted.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't. Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.

> Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.

Yes this is why I referred to 'months' and was downvoted by people who can't distinguish their politics from reality. You are restating exactly the text you are criticizing.

This is the nature of mechanical parrotlike repetition of propaganda:

a) You can have 'frontier' models with closed weights. Your sentence is basically a contraction within itself, again, the weight of ideology.

b) No one actually knows whether or not the closed weights of Bytedance are beyond all existing frontiers. Again the weight of ideology blinds you to the fact that the real 800 lb gorilla of Chinese Ai is more closed than Anthropic.

Re: DeepSeek-V4-Flash Update

#229

Earlier quoted context omitted.

That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold

Even Xi's speech was as paranoid as Dario Amodei about the possibility of Chinese AI achieving something at the level of say Mythos. If you seriously believe such a thing will be on Hugging Face I really don't know what to say.

There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity.

It's so good that the US government is rushing to ban all Chinese models as fast as they can.

Post reply on HN