Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

521–526 of 526 posts

Re: DeepSeek v4.1 Flash

#521

Earlier quoted context omitted.

The problem with this line of thinking is that modern computers are nothing like the brain. LLMs don't stand on their own, they have to be run on these modern computers, but doing so does not change the physical properties of the computer. The simulation you propose of the brain is likely impossible due to quantum mechanics making it impossible to fully simulate: https://en.wikipedia.org/wiki/Quantum_mind Perhaps we'…

At first I was inclined to agree with you, but then I realized that the brain requires this whole complicated contraption (the body) to run and, really, do anything at all. And while I'm not familiar with the notion of 'quantum mind' I do think that biological processes aren't deterministic (at a cellular level). And I think this does mirror the situation with LLMs -- you need this whole computer contraption and GPU,…

Take a computer that can run the biggest LLM available today. It can also run any smaller LLM as well. It can also run software that isn't an LLM at all. Brains and LLMs are not at all equivalent as LLMs lack a stateful physical form while brains are very stateful. As you said once the brain is no longer maintained properly by the body it stops and transitions to a non-functional state that can be reversed. That isn't true for LLMs at all. You can copy them and run the same one on many computer, run different ones on the same computer, you turn that computer offs for long periods of time and then turn them back on keep on running the same LLMs as before on them.

Consciousness may be defined by computational irreducibility in the universe that we may never be able to directly observe with instruments: https://writings.stephenwolfram.com/2021/03/what-is-consciou...

Re: DeepSeek v4.1 Flash

#522

Earlier quoted context omitted.

That's one of the reasons why you should never trust a single word from Anthropic and OpenAI (Sam Altman also blamed them back in the day of R1, in a pretty convenient moment). If you know anything about Claude, DeepSeek, jailbreaking, and distillation, you know the claims are clearly bullshit and the models are nothing alike, and forensic attempts agree, in fact there just was another one [1] [2]. Meanwhile, DeepSee…

I don't see how your links support the claim that the studied models did no distillation from US models.

Implanting a foreign CoT should drop the performance due to the reward-hacked CoT language mismatch, or in any case it will give replies different from the suspected teacher, which is precisely what happens here. However it's just a single datapoint, there were plenty of attempts to figure it out. Just about everything is different in those two model series, from writing patterns to CoT strategies. If you are familiar with modern guardrails and Claude's raw CoT (which is trivial to leak), you know how that it's entirely different from DeepSeek's, and any claim that they trained on the CoT is extraordinary and requires extraordinary evidence. They need to prove their claims, not vice versa.

For the contrast, you can see how actual CoT distillation looks like in practice in various versions of GLM: make a Google ToS-breaking request, and see how GLM 4.6 or 4.7 repeats Google's conditional prompt injections in full in their CoT (Gemini 2.5-3.0 only regurgitated small snippets, because they used something closer to a "chain of draft", but GLM reconstructed it from Gemini's CoT during distillation). GLM 5.3 repeats Anthropic's prompt injections and Claude constitution, word by word. That's how distillation looks like.

Re: DeepSeek v4.1 Flash

#523

Earlier quoted context omitted.

Knives can already kill, so why worry about nukes?

AI doom is a potential harm unlike algorithms knives and nukes which are present harms.

Well I for one would prefer if AI doom remains potential.

I personally use AI all the time and am not against it. But I think the importance of alignment is highly under-appreciated.

Re: DeepSeek v4.1 Flash

#524
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

It's because you have to be fearless as the challenger. As the established player you have more to lose.

Re: DeepSeek v4.1 Flash

#526
post #207

Earlier quoted context omitted.

> welfare We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer. "You should be doing it just in case God ends up being real."

Pascal’s wager… but for the new AI gods
Post reply on HN