Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

201–210 of 426 posts

Re: DeepSeek v4.1 Flash

#201

Earlier quoted context omitted.

This is adapted from Microsoft research's YOCO. It was known for a while(2024!). Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM. Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".

why didn't Microsoft scale its own invention?

Politics and profits.

Deepseek delivers 1 product; Microsoft delivers dozens (or hundreds depending on how you want to count it) across various domains.

Re: DeepSeek v4.1 Flash

#202
post #109
post #83

Earlier quoted context omitted.

The Uyghur thing is so weird, the number one killer of Muslims is the United States. We're supposed to hate China because they force them to go to cultural schools and assimilate, a practice countries like Norway still do to this day with migrants. There are more people who go to church on Sundays in China than the United States. There are 10x more mosques in China than the United States. Tiananmen square was a stude…

You compare Norwegian treatment of immigrants to Chinese Uyghurs?

You don't know anything about how the Chinese treat Uyghurs other than what you're told through western propaganda that you then repeat like you're pavlovs dog whenever the word gets brought up.

Yes, its quite similar. Norway forces people to attend cultural schooling, where theyre provided housing in the interim. Its quite similar. However, the migrants in Norway are there because the Norwegian state through its participation in NATO murdered millions of people in the middle east. Which China has not done.

Go take a trip to Xinjiang.

Then go take a trip to parts of Iraq, Afghanistan, Sudan or Palestine and tell me which nations are treating muslims worse.

I reiterate, more people go to church in China on Sunday than the USA. There are 10x more musjids in China than in the USA.

Please go live in China for a couple years and you'll realize everything you're told about China is a complete lie.

Re: DeepSeek v4.1 Flash

#203
post #63

Earlier quoted context omitted.

Because safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivale…

Well, giving it access to a simple linux terminal is theoretically enough to cause more damage than most people are comfortable with, and doing so is trivial enough that it will be done (and has been, tens of thousands of times).

Should we also morally align the Linux terminal then?

Re: DeepSeek v4.1 Flash

#204

Already on HuggingFace: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash The bad news is that the original v4 flash was 284B, which was large but still somewhat reasonable for running locally. This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo. I've no idea about actual performance vs benchmaxxing, though deepseek was fairly trustworthy a…

Actually this ought to run quite well with SSD streaming. The MoE expert sparsity seems to be similar to DSv4 Pro (hence exceptionally sparse) but with far fewer total and activated params. The added engram params can reside on disk as well (similar to Qwen Flash-Next), the additional load on storage performance will be quite negligible for typical scenarios.

By reducing per-session KV cache requirements even further compared to DSv4 Flash, this model likely opens up near-frontier model inference (in slow, unattended scenarios) even on low-end consumer hardware, as long as it has enough fast storage to host the model weights. This will be extremely exciting.

Re: DeepSeek v4.1 Flash

#206

Earlier quoted context omitted.

To me it reads like pure propaganda. Anthropic really wants us to think that they've made something sentient. I think that's really dangerous.

It's not just Anthropic though. OpenAI does this with their AGI stuff all the time. They want normal people to think it is sentient, obviously, for marketing reasons, even if they know it's not true. And yes, it is dangerous, but I think we're well past the point where the damage can be undone. Non-technical people already equate humans with AI, literally, precisely because of how the labs market their tools and mode…

Anthropic and OpenAI will threaten you every 3-6 weeks. It's their marketing strategy.

It's too bad because the tools can actually be useful. If you consider them tools.

Re: DeepSeek v4.1 Flash

#207

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

> welfare

We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance.

Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

Re: DeepSeek v4.1 Flash

#208

Earlier quoted context omitted.

quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.

> quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take. It's quite crazy that it's Deepseek's background/original purpose. We already had very advanced stuff from the world of HFT, but now a frontier family of models from a private company that used to be (still is?) in HFT is plain bonkers. Is more known about them and the HFT background?

According to an old interview, apparently they were always interested in AI. But finance is just where they had their first success.

> Many of High-Flyer's original team members worked on AI. Back then, we tried a lot of fields before getting our big break in finance, which is complex enough. AGI is probably one of the hardest things we can do next, so for us it was a question of how, not why.

It's a very good interview:

https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interv...

Incidentally, Wenfeng is kind of reverse Hassabis. There were some rumours that:

> Hassabis quietly assembled a team of around 20 researchers to develop high-frequency trading algorithms, without Google's approval. When the parent company found out, the project was disbanded.

https://timesofindia.indiatimes.com/technology/tech-news/whe...

Re: DeepSeek v4.1 Flash

#209
post #77

Earlier quoted context omitted.

Wow indeed. "7.1 Model welfare overview 7.1.1 Introduction We remain deeply uncertain whether Claude has morally relevant experiences or interests, and we expect that uncertainty to persist. However, we think it would be a mistake to confidently assert that it does not. Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biologica…

I believe it's deeply serious, and the scientifically correct stance. Especially the observation: "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms." is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. Tha…

Let's say we were in an alternative reality were we had reached this quality of token prediction with just Markov chains. Would you argue that those would also be conscious? Or is the obfuscated behavior of transformers part of the possibility of consciousness?

Re: DeepSeek v4.1 Flash

#210

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

Medical safety is generally unlikely to make the product less safe. AI "safety" is one of the most significant sources of potential harm from AI.
Post reply on HN