If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.
DeepSeek v4.1 Flash
21–30 of 411 posts
Re: DeepSeek v4.1 Flash
#22I was hoping for a bit more, but it's still 100% faster for a very good price, so I won't complain.
Re: DeepSeek v4.1 Flash
#23Every model release seems like it packed with wonderful research and advancements.
Re: DeepSeek v4.1 Flash
#24Re: DeepSeek v4.1 Flash
#25Re: DeepSeek v4.1 Flash
#26Already on HuggingFace: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash The bad news is that the original v4 flash was 284B, which was large but still somewhat reasonable for running locally. This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo. I've no idea about actual performance vs benchmaxxing, though deepseek was fairly trustworthy a…
V4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB. Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines. Edit: Most of added weights/size are Engrams? > Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.…
552B in ~FP4, 306GB.
196B of FP8 Engrams, another 204GB, not necessary to keep in RAM.
KV cache sees another 4x size reduction, just 900MB for 1M.
So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.Re: DeepSeek v4.1 Flash
#27It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
Re: DeepSeek v4.1 Flash
#28Re: DeepSeek v4.1 Flash
#29If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.
I'm starting to have chinese characters bleed into claude as well. Perhaps a sign of the times. Understanable for a chinese first model but an english first (supposedly) model? wild stuff.
Re: DeepSeek v4.1 Flash
#30It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
> Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
Do the Chinese have no such scientists?