Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

91–100 of 424 posts

Re: DeepSeek v4.1 Flash

#91
post #70

While this is very impressive benchmark-wise, GPT-6 Astra showed us that benchmarks don't always correlate 1:1 to intelligence of a model. When Astra launched, I think Artifical Analysis showed that it was on par with GPT-5.6 Sol and lower than Opus or something like that? Then, they updated the scoring. I hope that more open source models, including this model, to be "as good to use" as Astra.

I don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.

Re: DeepSeek v4.1 Flash

#92
post #26
post #6

Earlier quoted context omitted.

V4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB. Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines. Edit: Most of added weights/size are Engrams? > Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.…

It's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x size reduction, just 900MB for 1M. So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.

Or two gorgon halos?

Re: DeepSeek v4.1 Flash

#93
post #16

I'm a big fan of DeepSeek. Also, ask it what model it is :) In Pi (pi.dev), it tells me it's definitely Claude by Anthropic, via the API via curl it tells me it's "probably ChatGPT", its very funny.

Works correctly in opencode, but seems like they inject a system prompt:

Thinking: > The user is asking what model I am. According to my system prompt, I'm powered by "deepseek-flash" with model ID "opencode-go/deepseek-flash".

>I'm powered by the model opencode-go/deepseek-flash.

Re: DeepSeek v4.1 Flash

#95

My question is: what kind of hardware do you need to run this Flash beast locally at a meaningful speed?

8x RTX PRO 6000 or 4x Spark? Or 1x M5 Ultra 512GB.

The model is theoretically FP8, but really internally its mostly FP4 already, so there won't be a cut-in-half-but-almost-just-as-good quant coming for this one.

Re: DeepSeek v4.1 Flash

#97
post #47
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

[flagged]

Google's Gemma models are usually celebrated, so were the llamas. If Meta releases Muse Spark it will also be a good thing. If Anthropic released a great open weight model I am sure that post won't be steered towards controversy and anti-AI sentiment.

It so happens Chinese companies are more friendly towards open weights, autonomy and freedom that most US based ones. Who would have guessed?

Re: DeepSeek v4.1 Flash

#98
post #84

Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristicti…

I do wonder how long this will last. I bet in a few month or years they all have similar ~legal~ blocks.

Great thing about it, since it's open weight those blocks can easily be ablitared away

Re: DeepSeek v4.1 Flash

#99
post #27

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

Anthropic's stance on safety it's just PR management and their hope to keep the others down, they are rushing as blind as everyone else to whatever improvement they can achieve.

Re: DeepSeek v4.1 Flash

#100
post #39

Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristicti…

> My favourite benchmark for this is to ask it to download a rom for an old game Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked. And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.

Or if you apply to a company and they want to do an AI HR interview and an AI coding test and an AI challenge - if you throw OpenAI or Claude models at it - they refuse, because it's "wrong" and "immoral".

Not so with the Chinese models.

Post reply on HN