Earlier quoted context omitted.
Not yet, right? That's the old DS V4 preview release. We're still waiting for the weights to come out.
Probably the same. When the same base model is trained, the weight do not tend to change a lot. GLM 5.0 > 5.1 > 5.2 are the same base model, that just kept being trained. Weights hardly change as a result. Think in the like few percentage points size difference.
DeepSeek-V4-Flash Update
211–220 of 362 posts
Re: DeepSeek-V4-Flash Update
#212Earlier quoted context omitted.
Do you have any recommendations of such extensions for pi?
Recommendation? No. Just go with the passive-aggressive advice "let pi build it for you". :-) To be more constructive, what I did (as an experiencd SWE but a complete noob to agentic coding): went to pi.dev's extension marketplace and looked into all the new shiny stuff. Subagents, mcps, context and memory optimizers, skills. Using the most popular ones (not necessarily the best ones) It was like 15years ago learning…
Re: DeepSeek-V4-Flash Update
#213Earlier quoted context omitted.
I agree a bio is a bad proxy, and I'm not saying I'm amazing but I'm not THAT incompetent. Most of my GitHub is pre-LLM era; github.com/lionkor Edit: And like a lot of tools, HOW you use them is just about as important as the quality of the tool itself. Of course the tools can produce massive amounts of bad quality slop, they can also produce fast, focused edits that make sense.
I'm not even trying to call you out, I'm sure you're solid.
Re: DeepSeek-V4-Flash Update
#214Re: DeepSeek-V4-Flash Update
#215Deepseek and moonshot are the only two providers I consent to training for.
Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.
Re: DeepSeek-V4-Flash Update
#216Earlier quoted context omitted.
Unfortunately, all mentions of ZDR have silently been removed from the OpenCode Go page today.
Thanks, any update here is important. I use them because of good data policies.. I still find this today: > The plan is designed primarily for international users and provides stable global access. Your data will not be used for model training.
Re: DeepSeek-V4-Flash Update
#217What's the best way to run this on a 64GB M2 Pro?
Re: DeepSeek-V4-Flash Update
#218In case people want to run it, it's DeepSeek-V4-Flash-284B-A13B. So it should just barely run on a single B300, and it's small enough that it'll barely run on an M5 Max too.
It runs really well on 2 DGX Sparks - 60t/s
Dual Strix Halo is much slower and current Macs with 256GB are both slower and more expensive (Mac Studio M3 Ultra 256GB around 12000€).
To get something faster than the two Sparks you'd need to spend more than $22000 for a server with 2x RTX Pro 6000 at $10000 each.
Beyond that you could get 2x AMD MI350P.
Re: DeepSeek-V4-Flash Update
#219Earlier quoted context omitted.
Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.
What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't. Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.
Re: DeepSeek-V4-Flash Update
#220Deepseek and moonshot are the only two providers I consent to training for.
Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.