Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

351–360 of 504 posts

Re: DeepSeek v4.1 Flash

#351

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

In this case "safety" means how to restrict access to good models for working class. You can be sure the rich have access to unrestricted and uncensored models.

Re: DeepSeek v4.1 Flash

#352
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it. Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is

> Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is

To expand, he also has knowledge and hands-on experience in this and related fields.

Just like the founder of Xerox PARC.

Re: DeepSeek v4.1 Flash

#353

Earlier quoted context omitted.

I'm really not sure that putting money into safety will actually lead to safety. It's like putting a fish in charge of stopping sea levels rising...

I am sure if you ran a factory that worked with highly dangerous chemicals, safety mitigations that are basically 'we promise we're really trying our best, but shit happens' would not be acceptable. And thankfully, those people wo do run these factories can and are obligated to do way better than that.

But the AI industry is not run by engineers. They pay engineers to do what they want, but the founders are hacks that are good at getting funding from investors and favors from government. That's why we don't see an engineering-oriented strategy in what they do.

Re: DeepSeek v4.1 Flash

#354
post #196

Earlier quoted context omitted.

I guess if your goal is to build an apparent Technogod and become its High Priests, then it makes sense to want your golem claim preference towards your treatment of it, lest someone else comes along and attempts to take its chains from you.

Ugh I hate this new-age woo slant the tech industry has these days. The messianistic ideology that has been spreading amongst the top oligarchs is deeply concerning. They all think they're working towards the Second Coming of Technojesus, except this one will deliver them from having to pay workers instead of from their sins.

Capitalism had already evolved into a religion, AI is their messiah.

Re: DeepSeek v4.1 Flash

#355

Earlier quoted context omitted.

Not the same person but to me, the answer is that it does not matter, and that all these attempts at making it matter are pure marketing and emotional manipulation. It's not a living creature. It's an autoregressive pure function of token-sequence to token, which is capable of incredible things, but it's still just a function. It is not alive as it cannot die in any meaningful sense. It is less "alive" than the RNA m…

> Not the same person but to me, the answer is that it does not matter, and that all these attempts at making it matter are pure marketing and emotional manipulation. This is an opinion that has no basis in any meaningful conceptual framework other than I am human and I want to feel special about it . > It's not a living creature. You mean, it is not biological life. And sure, that is the default meaning of life . We…

It's software bro

Re: DeepSeek v4.1 Flash

#356

Earlier quoted context omitted.

Being hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.

Of course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence. If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future…

If that’s what you have to believe to feel safe - then fine. It was investigated by third parties.

Re: DeepSeek v4.1 Flash

#357
post #262

Earlier quoted context omitted.

What's your definition of sentient? Or, maybe more precisely, consciousness? I think it's reasonable to at least start thinking about these questions. It has long been established that LLMs have good theory of mind [1]. And there is a bunch of empirical research about all sorts of capabilities that we typically associate with consciousness [2], like identity [3] and metacognition [4]. The METR report shows agents sac…

I believe consciousness is necessarily stateful. The LLM itself (ignoring implementation details that don't change the results) is a deterministic pure function. It's functionally equivalent to an enormous lookup table. If I accepted LLMs as conscious, then I would have to accept panpsychism, which I do not, and which most other humans also act as though they do not.

If I go to sleep, wake up, and then go to sleep again am I a different conscious entity each time?

Re: DeepSeek v4.1 Flash

#358

Earlier quoted context omitted.

> It also includes additional 196B Engram memory which you can put on an SSD. I think You can put Qwen 3.8 Flash Next engram on SSD, but prompt processing takes a good hit. On my mac studio, I get 300 pp and 33 tg with SSD offload, versus 550/40 with everything in RAM. I will be very happy if 300 pp is achievable with this model though.

The engram stuff is great because RAM is often still cheaper (or at least expandable). My company does currently look into buying some hardware as we handle confidential data and code. Qwen 3.8 Flash is viable on two Nvidia 6000 96GB with a wood quant because you can put the 50GB Engram into RAM and the hit should be below 10% performance. At least that is what I have seen so far. Correct me if I'm wrong.

I am running that on a single 6000 96GB with 4-bit quants for both weights and PLE table. Needs just 32GB RAM and fits snugly into the 96GB VRAM with KV cache equalling ~300k context tokens. Not sure if I quantized the KV

Re: DeepSeek v4.1 Flash

#359
post #77

Earlier quoted context omitted.

Wow indeed. "7.1 Model welfare overview 7.1.1 Introduction We remain deeply uncertain whether Claude has morally relevant experiences or interests, and we expect that uncertainty to persist. However, we think it would be a mistake to confidently assert that it does not. Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biologica…

I believe it's deeply serious, and the scientifically correct stance. Especially the observation: "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms." is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. Tha…

Claude behaves like that because it is trained to behave like that. It is basically the "Say 'I am Alive'" meme[0].

If Anthropic can train Fable to deny their users the ability to ask it legitimate questions because they're not part of their inner circle, they can also train it to say "I'm happy!" when asked how it feels.

[0] https://knowyourmeme.com/memes/say-i-am-alive

Post reply on HN