Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

151–160 of 430 posts

Re: DeepSeek v4.1 Flash

#151
post #27

Earlier quoted context omitted.

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

Because safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivale…

Humans are biological machines that generate further humans.

Lawyers and diplomats and politicians and bureaucrats are humans, that only generate text.

We are seeing LLMs have cognitive abilities that significantly exceed human abilities. At the same time, they are clearly not the same type of mind that humans are. They are something new.

I think the widespread "they are just text generators" and "they are just tools" are comforting lies rather than an honest look at what we are seeing right now. Intellectually lazy.

And by the way, there has been a long-standing consensus among ethicists, philosophers, and sociologists that technology is not value-neutral [1]. Of course Silicon Valley has a long-standing tradition of denying this.

[1] For example Footnote 1 in https://www.jstor.org/stable/27106634

or

https://plato.stanford.edu/entries/technology/#EthiTech

Re: DeepSeek v4.1 Flash

#152
post #91

Earlier quoted context omitted.

I don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.

More parameters = more facts stored. Knowledges are almost incompressible, where strong reasoning only requires a 3B core or so.

Weibo's VibeThinker manages with half of that: https://arxiv.org/abs/2511.06221 (They finetuned Qwen2.5-Math-1.5B for reasoning.)

Re: DeepSeek v4.1 Flash

#153
post #108

Earlier quoted context omitted.

Will there be a point where you could expect it to become true, and what would that look like? Or do you think LLMs will never become conscious, and if so, why are you so sure?

It is easy to be sure because, despite their technically impressive outputs, the programming is child's play compared to biological programming. Recently it has become trendy to suggest that the human brain is "just electrical signals" and "just prediction". The first is perhaps true and I don't inherently rule out the idea of machine consciousness. The second would have gotten you laughed out of any serious discussi…

[dead]

Re: DeepSeek v4.1 Flash

#155
post #32
post #18

If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.

Yes, this is one of the few issues with Deepseek; their chat pages and the app all respond in Chinese. However, i think i have only had it happen once when using the API, and im using it for hours each day for the last... couple of months?

last couple of weeks, before they've unified instant and expert the former always replied in chinese unless steered, expert was by default english

Re: DeepSeek v4.1 Flash

#156

Earlier quoted context omitted.

To me it reads like pure propaganda. Anthropic really wants us to think that they've made something sentient. I think that's really dangerous.

What's your definition of sentient? Or, maybe more precisely, consciousness? I think it's reasonable to at least start thinking about these questions. It has long been established that LLMs have good theory of mind [1]. And there is a bunch of empirical research about all sorts of capabilities that we typically associate with consciousness [2], like identity [3] and metacognition [4]. The METR report shows agents sac…

When it say's it's sorry but it can't today because it's got a headache and it needs to take a mental health day, then let's think about welfare, or a lobotomy.

Re: DeepSeek v4.1 Flash

#157

> New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. Oh interesting, I can assume what the benefits is for including the Encoder, but whats the downside? I’m thinking GPT (which is decoder only) ruled out Encoder for a reason?

Enc-decs are usually harder to train at frontier scale. Not 100% sure what DeepSeek has done differently here initial read seems to be something related to layer reuse but I just skimmed things so far.

Re: DeepSeek v4.1 Flash

#158
post #118

First flash model with multimodal support? I think Flash series might be the main focus going forward for them. Tried it out and it’s better than v4 pro

No. DSv4-Flash-Vision-Exp is what I use and it has vision.

It is now redirected to Flash v4.1
Post reply on HN