I think it's very clear that DeepSeek is obviously the best AI lab in the world. Every model release seems like it packed with wonderful research and advancements.
DeepSeek v4.1 Flash
271–280 of 408 posts
Re: DeepSeek v4.1 Flash
#272Re: DeepSeek v4.1 Flash
#273Earlier quoted context omitted.
LLMs may be conceptually simple, simpler than human brains but I don't see how that would prove that they cannot be conscious. Complex behavior can emerge from very simple rules. I would disagree that they aren't improving on obvious failure modes, but what does it even prove anyway? We know their intelligence is different than from a human, that doesn't mean they cannot be conscious. Would you change your mind if th…
> Complex behavior can emerge from very simple rules. Indeed. You can observe emergent behaviour from, for instance, Conway's Game of Life, written in 1970. Redefining consciousness as "has emergent behaviour" is another take that would have rightfully gotten one ridiculed 5 years ago. > but I am not certain and I don't see a way to be certain. One way to be certain is to reason about it. They are programmed to do no…
There's no print statements or human entered logic involved in the raw model expression at all.
The only thing that humans have programmed is efficient parallel dot product pipelines that "animate" (for lack of a better word) the models.
Everything these models do is emergent from their backpropgation guided evolution. This even includes in context learning itself, which was not an expected outcome.
Re: DeepSeek v4.1 Flash
#274Earlier quoted context omitted.
A stab: a video recording of a biological organism can exhibit many markers that would indicate consciousness if observed in a biological organism.
A video is a fixed representation. What if we can interact with this video, and it reacts in the same ways the source organism does? Then we put it in new situations that weren't in the source video, and it interacts in a similar way to the original organism in these situations, too. What do we make of reactions of pain or joy? Where's the line between simulation and enaction? This is closer to the reality of these m…
AFAIK using the same input tokens, weights, and numerical operations will lead to the same probability distribution for the next token. It uses pseudo-randomness to enable temperature, etc. Like a fuzzy video.
"Markers that would indicate consciousness if observed in a biological organism" just does not mean very much. A PR phrase used to hype the IPO.
Re: DeepSeek v4.1 Flash
#275Earlier quoted context omitted.
It's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x size reduction, just 900MB for 1M. So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.
Or two gorgon halos?
Re: DeepSeek v4.1 Flash
#276I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.
I'm really not sure that putting money into safety will actually lead to safety. It's like putting a fish in charge of stopping sea levels rising...
And thankfully, those people wo do run these factories can and are obligated to do way better than that.
Re: DeepSeek v4.1 Flash
#277Absolutely insane performance and benchmark results. It's beating Opus 5 and Sol 5.6 https://tokenstead.ai/models/deepseek-v4-1-flash
From the pricing, it's 3x cheaper on cache, 1/3 more expensive on input, and equal on output compared to GPT 5.6 Luna.
I would love to compare these two at work, where I pay API prices.
At home I will stick to Astra and Fable.
Re: DeepSeek v4.1 Flash
#278Earlier quoted context omitted.
Because we've been told these models are too dangerous since GPT2. At this point it's just marketing stunts.
> At this point it's just marketing stunts. If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing. It seems…
When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.
Re: DeepSeek v4.1 Flash
#279Earlier quoted context omitted.
> Much more knowledgeable than you or I are Speak for yourself. I work for an LLM startup that was successfully bootstrapped and is now highly profitable with 8-digit revenue and zero outside investment. Unlike OpenAI and Anthropic, we do not rely on deceiving investors to dump a trillion dollars into a tar fire with the false promise of delivering the machine god that will unemploy all of humanity (at best). Taking…
> Speak for yourself. I work for an LLM startup And yet you still fail to demonstrate good understanding of the topic ¯\_(ツ)_/¯ > stating that those are the only people who can be trusted You are right, they are most definitely not the only people who can be trusted to have current and accurate information. But due to the unique constraints of these fast-moving events, they are certainly among those whose opinions ne…
Or you simply misinterpreted my words, seemingly intentionally so because pedantry is a comfortable fall-back for not having a logical argument.
> Saying (derisively) that it is an "appeal to emergent behaviour", when the ENTIRE POINT OF CONTENTION is said emergent behaviour is like saying that you should not discuss God at a theological forum or that you should ignore the theory of relativity when discussing gravity.
The derisiveness comes from the fact that you appear to believe merely demonstrating emergent behaviour is enough, despite the fact that emergent behaviour is common and has been common in programs for half a century without anybody considering them conscious. Life itself is emergent behaviour, but that does not mean all emergent behaviour is life. Life emerged from incredibly complex physical and material interactions over billions of years of incremental self-programming. The idea that we have found some magic ingredient to shortcut the process, that we can recreate that with some very simple statistical model that is not capable of self-programming, is so absurd it becomes about as difficult to argue against as Russell's Teapot. We developed a model for predicting words and it does. Although it does quite an impressive job of that, it has demonstrated zero capability to do anything beyond what you would reasonably expect it to, same as all other software with emergent capabilities and rather unlike life which developed truly novel emergent behaviour relative to its base ingredients.
Re: DeepSeek v4.1 Flash
#280Earlier quoted context omitted.
They have flat fees, so it's the best deal around by far. Basically for $5 first month then $10/mo after that. If you're doing tons of heavy work, it struggles because they throttle the model inference and for good reason. I mean it's cheap! But if you want a place to try models for nearly nothing and aren't doing 6 sessions in parallel it works fine.
How many tokens are you getting, roughly?
For deepseek-v4-flash: a shitton of tokens for $5.