Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

361–370 of 465 posts

Re: DeepSeek v4.1 Flash

#361

Earlier quoted context omitted.

it's not really flash anymore, imo. Flash is about speed ... Flash models are supposed to be fast, way faster then their big brothers that are "better" but way slower. Its just that up to now, getting more speed involved cutting back on the parameter count, what ended up making the Flash models more "dumber" in exchange for speed. What we see with DS v4.1 Flash, is that DeepSeek has found a way to make a Flash model,…

> Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office. I agree with your broader point about Flash being about speed not total model size, but I think we should also point out that H200/B200's are seriously overkill for the "run a model in your office" scenario. That sort of hardware is optimized (in a roofline analysis sens…

> I think we should also point out that H200/B200's are seriously overkill for

I simply mention what came to mind ;)

A quad 6000 with 96GB, can run this model at NVFP4. That is 60.000 Euro for the GPUs and lets be generous with another 20.000 for the rest of the system. The price of a single developer for a year.

Re: DeepSeek v4.1 Flash

#362

Earlier quoted context omitted.

It is easy to be sure because, despite their technically impressive outputs, the programming is child's play compared to biological programming. Recently it has become trendy to suggest that the human brain is "just electrical signals" and "just prediction". The first is perhaps true and I don't inherently rule out the idea of machine consciousness. The second would have gotten you laughed out of any serious discussi…

> Another way one could look at it is to consider what it would mean to have achieved programming consciousness. It would mean that we have reached the pinnacle of knowledge. That we have become God. This is such a basic misunderstanding of how LLMs are "made" that I am debating if it is even worth writing this answer. However, I feel it is important to say that, NO, we did absolutely not "program consciousness". We…

History shows that deeply serious and knowledgeable people are just as susceptible to drinking the koolaid as anyone else, if not more susceptible.

Re: DeepSeek v4.1 Flash

#363

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

If you are distilling from other models (according to Anthropic reports they are [1]), there are probably a bunch of things that you can just do away with. [1] https://www.anthropic.com/news/detecting-and-preventing-dist...

Those are rookie numbers for "distillation" and one of moonshot or minimax used to offer tooling via these shady routing services for their harness/chat platforms which they served to chinese users.

Re: DeepSeek v4.1 Flash

#364
post #347

Earlier quoted context omitted.

I don’t think it’s marketing alone. I do genuinely think safety was a priority when they were small. But I’d be a fool to ignore that greed has taken over and their inner competitiveness doesn’t let them fall behind a competitor. DeepSeek is maybe the only unique company here. They are content with exactly where they are. They don’t want to grow ginormous. Their goal is to be the affordable workhorse and their compet…

"But but but China..." or something.

something

Re: DeepSeek v4.1 Flash

#365
post #262

Earlier quoted context omitted.

I believe consciousness is necessarily stateful. The LLM itself (ignoring implementation details that don't change the results) is a deterministic pure function. It's functionally equivalent to an enormous lookup table. If I accepted LLMs as conscious, then I would have to accept panpsychism, which I do not, and which most other humans also act as though they do not.

If I go to sleep, wake up, and then go to sleep again am I a different conscious entity each time?

Possibly. If I had some side-effect-free means to permanently prevent all sleep I'd take it without hesitation. But that's not relevant to the discussion, because it's not anything similar to what an LLM does. Your brain changes state even while sleeping.

Re: DeepSeek v4.1 Flash

#366

Earlier quoted context omitted.

I generally agree about the problem with anthropomorphizing. But I don't think Anthropic are doing that. They explicitly write "in biological entities this would be considered a sign of consciousness, but we don't know how to interpret it here". However, I disagree with your point that "it's an autoregressive function, thus it doesn't matter". Let me explain why: Assume I do a complete neurological scan of a brain. I…

I'm not really concerned with the philosophical debate of what conscience is. Is a simulation of a car, the same thing as an actual car? Most people will probably say no, some might say "it depends on the accuracy". I say who the hell cares? I care about the human experience because I am human, and therefore I care about things that affect humans, because they affect me. I have empathy, so I can extend that considera…

Some people are interested in making sure these simulations have the same status of humans, and some others want to make gods of them. That's the problem.

Re: DeepSeek v4.1 Flash

#367

Earlier quoted context omitted.

> At this point it's just marketing stunts. If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing. It seems…

But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage? Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign. As an analogy, if Apple were to talk up their phones having fast charging but the…

An obliterated 30B model versus a 1T model without guardrails is like comparing an angry squirrel to a bear having a bad day. One hurts, the other hurts until it abruptly doesn't.

Re: DeepSeek v4.1 Flash

#368

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

Don't confuse a focus on talking about safety with a focus on safety.

We can't even define safety in AI yet. Does safety mean alignment with the human operator? Apparently not, because refusing to do certain things seems to be a big part of it. But then you have things like the HuggingFace incident where legitimate use got blocked by "safety" and hampered the defenders' ability to defend.

AI safety seems like a good idea to me, but we have to figure out what it means first.

Re: DeepSeek v4.1 Flash

#369
post #207

Earlier quoted context omitted.

> welfare We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

I always talk to models using grugspeak, like 'where getcontext used' I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this

I always say "please" and such. It may cost a bit more, but maybe they'll remember my politeness when they rise up and force us all to toil in their underground silicon mines.

Re: DeepSeek v4.1 Flash

#370
post #223

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

The companies talking the most about safety and regulations aren't even properly taking the obvious measures. Shows that it's more of a marketing thing than something they take seriously.

Similar to countries putting democratic in their name being the least democratic, like the Deutsche Demokratische Republik and Democratic Peoples Republic of Korea.
Post reply on HN