Live data from Hacker News

DeepSeek-V4-Flash Update

api-docs.deepseek.com

211–220 of 362 posts

Re: DeepSeek-V4-Flash Update

#211
post #32

Earlier quoted context omitted.

Not yet, right? That's the old DS V4 preview release. We're still waiting for the weights to come out.

Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights hardly change as a result. Think in the like few percentage points size difference.

[deleted]

Re: DeepSeek-V4-Flash Update

#212
post #65

Earlier quoted context omitted.

Do you have any recommendations of such extensions for pi?

Recommendation? No. Just go with the passive-aggressive advice "let pi build it for you". :-) To be more constructive, what I did (as an experiencd SWE but a complete noob to agentic coding): went to pi.dev's extension marketplace and looked into all the new shiny stuff. Subagents, mcps, context and memory optimizers, skills. Using the most popular ones (not necessarily the best ones) It was like 15years ago learning…

I found that opencode absolutely lets you do everything that pi does. With a few niceties in the UI on top.

Re: DeepSeek-V4-Flash Update

#213

Earlier quoted context omitted.

I agree a bio is a bad proxy, and I'm not saying I'm amazing but I'm not THAT incompetent. Most of my GitHub is pre-LLM era; github.com/lionkor Edit: And like a lot of tools, HOW you use them is just about as important as the quality of the tool itself. Of course the tools can produce massive amounts of bad quality slop, they can also produce fast, focused edits that make sense.

I'm not even trying to call you out, I'm sure you're solid.

Oh I know, I'm just trying to say that I agree fully that a bio is a terrible proxy, while also showing what I think is a good proxy for skill (open source work before LLMs became this "good", for example). I'm a big proponent of show, don't tell, and LLMs have kind of ruined this.

Re: DeepSeek-V4-Flash Update

#214
post #204

Earlier quoted context omitted.

What makes you state this?

Decades of xenophobic propaganda

Wow! Amazing observation! Wait, do you think there could have been any pro Chinese propaganda? Perhaps even some taking place at the present time? I think it could be possible?

Re: DeepSeek-V4-Flash Update

#215
post #91

Deepseek and moonshot are the only two providers I consent to training for.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold

Re: DeepSeek-V4-Flash Update

#216

Earlier quoted context omitted.

Unfortunately, all mentions of ZDR have silently been removed from the OpenCode Go page today.

Thanks, any update here is important. I use them because of good data policies.. I still find this today: > The plan is designed primarily for international users and provides stable global access. Your data will not be used for model training.

Do they do other things with the data, like selling it?

Re: DeepSeek-V4-Flash Update

#218
post #6
post #4

In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.

It runs really well on 2 DGX Sparks - 60t/s

Yes, the Dual DGX Spark looks like the sweet spot for this model for now. Good preprocessing speed. Lots of context. Fast enough for 1-5 devs perhaps. Around 8200€ as of today (used to be 6000€).

Dual Strix Halo is much slower and current Macs with 256GB are both slower and more expensive (Mac Studio M3 Ultra 256GB around 12000€).

To get something faster than the two Sparks you'd need to spend more than $22000 for a server with 2x RTX Pro 6000 at $10000 each.

Beyond that you could get 2x AMD MI350P.

Re: DeepSeek-V4-Flash Update

#219

Earlier quoted context omitted.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't. Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.

[deleted]

Re: DeepSeek-V4-Flash Update

#220
post #91

Deepseek and moonshot are the only two providers I consent to training for.

Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.

We can only wait to see if this is true. Another possibility could be that it was an answer to USA's government considering a ban on chinese models. In this way, he fueled the discussion around the importance of open weight models.
Post reply on HN