While this is very impressive benchmark-wise, GPT-6 Astra showed us that benchmarks don't always correlate 1:1 to intelligence of a model. When Astra launched, I think Artifical Analysis showed that it was on par with GPT-5.6 Sol and lower than Opus or something like that? Then, they updated the scoring. I hope that more open source models, including this model, to be "as good to use" as Astra.
DeepSeek v4.1 Flash
91–100 of 474 posts
Re: DeepSeek v4.1 Flash
#92Earlier quoted context omitted.
V4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB. Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines. Edit: Most of added weights/size are Engrams? > Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.…
It's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x size reduction, just 900MB for 1M. So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.
Re: DeepSeek v4.1 Flash
#93I'm a big fan of DeepSeek. Also, ask it what model it is :) In Pi (pi.dev), it tells me it's definitely Claude by Anthropic, via the API via curl it tells me it's "probably ChatGPT", its very funny.
Thinking: > The user is asking what model I am. According to my system prompt, I'm powered by "deepseek-flash" with model ID "opencode-go/deepseek-flash".
>I'm powered by the model opencode-go/deepseek-flash.
Re: DeepSeek v4.1 Flash
#94Significant jump in pricing. V4 Flash was $0.16/M out, 4.1 is $1.20/M.
Re: DeepSeek v4.1 Flash
#95My question is: what kind of hardware do you need to run this Flash beast locally at a meaningful speed?
The model is theoretically FP8, but really internally its mostly FP4 already, so there won't be a cut-in-half-but-almost-just-as-good quant coming for this one.
Re: DeepSeek v4.1 Flash
#96Re: DeepSeek v4.1 Flash
#97As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…
[flagged]
It so happens Chinese companies are more friendly towards open weights, autonomy and freedom that most US based ones. Who would have guessed?
Re: DeepSeek v4.1 Flash
#98Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristicti…
I do wonder how long this will last. I bet in a few month or years they all have similar ~legal~ blocks.
Re: DeepSeek v4.1 Flash
#99It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
Re: DeepSeek v4.1 Flash
#100Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristicti…
> My favourite benchmark for this is to ask it to download a rom for an old game Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked. And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.
Not so with the Chinese models.