Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

331–340 of 411 posts

Re: DeepSeek v4.1 Flash

#331

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

Wow there really is a model welfare section in there...

My fault I guess, verbally abusing Claude in my experience gives better results.

Re: DeepSeek v4.1 Flash

#333

Earlier quoted context omitted.

I'm not really concerned with the philosophical debate of what conscience is. Is a simulation of a car, the same thing as an actual car? Most people will probably say no, some might say "it depends on the accuracy". I say who the hell cares? I care about the human experience because I am human, and therefore I care about things that affect humans, because they affect me. I have empathy, so I can extend that considera…

Your car analogy is very, very confused. If a simulation of a car can get you from A to B, requires the same steering, fuel and servicing, gives you the same tactile feedback, is it a car? Of course "I care about humans because I am human" is a self-consistent position to take. But now you need to decide if you want to consider a full simulation that faithfully reproduces everything that physically happens between ou…

> If a simulation of a car can get you from A to B, requires the same steering, fuel and servicing, gives you the same tactile feedback, is it a car?

No, because a simulation of a car cannot get me from A to B. No matter how accurate you make it, I can't get to my supermarket with it, because it's just a bunch of math on a computer.

It's an interesting sort of self-defeating position, the whole "simulation of human consciousness = human consciousness", because it simultaneously attempts to devalue the human experience, while also elevating the importance of a particular human brain process.

A robot running a simulation of the human mind is a robot, not a human.

Re: DeepSeek v4.1 Flash

#334

Absolutely insane performance and benchmark results. It's beating Opus 5 and Sol 5.6 https://tokenstead.ai/models/deepseek-v4-1-flash

If it's actually comparable in practice that would be very impressive. I am yet to try a DeepSeek model. From the pricing, it's 3x cheaper on cache, 1/3 more expensive on input, and equal on output compared to GPT 5.6 Luna. I would love to compare these two at work, where I pay API prices. At home I will stick to Astra and Fable.

Just out of curiosity, may I know why you pay API price at work versus using dozens of subscriptions that you APIfy?

Re: DeepSeek v4.1 Flash

#335
post #77

Earlier quoted context omitted.

Wow indeed. "7.1 Model welfare overview 7.1.1 Introduction We remain deeply uncertain whether Claude has morally relevant experiences or interests, and we expect that uncertainty to persist. However, we think it would be a mistake to confidently assert that it does not. Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biologica…

I believe it's deeply serious, and the scientifically correct stance. Especially the observation: "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms." is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. Tha…

I can feed my biological markers into a set transformer with the time of day, what I'm doing, what I ate, if I'm on-call, and it'll predict my next glucose, heart rate, blood pressure, melatonin, etc state quite well. It's still just a transformer without hormones, blood vessels or glucose metabolism, no matter how well it internally represents metabolic distress markers.

Re: DeepSeek v4.1 Flash

#336

Earlier quoted context omitted.

Being hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.

Of course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence. If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future…

[dead]

Re: DeepSeek v4.1 Flash

#338

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

Wow there really is a model welfare section in there...

https://icml.cc/virtual/2026/poster/67058

Re: DeepSeek v4.1 Flash

#340
post #223

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

The companies talking the most about safety and regulations aren't even properly taking the obvious measures. Shows that it's more of a marketing thing than something they take seriously.

If I operate a nuclear reactor or a hydroelectric dam there are regulators that tell me what i'm allowed to do, so as to keep my profit motive from overwhelming the public interest.

If we want AI to actually have some safety rails, this is what we would do.

Post reply on HN