Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

161–170 of 478 posts

Re: DeepSeek v4.1 Flash

#161
I've run some evals on my puzzle game https://redactle.net/llm-leaderboard

Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest.

I'm curious what other unique evals people are running.

Re: DeepSeek v4.1 Flash

#162

Earlier quoted context omitted.

Hit me up if anyone wants extra $5 free usage with my referral code

Never tried OpenCode Go so I am interested. How does their pricing compare to paying DeepSeek directly, by the way?

They have flat fees, so it's the best deal around by far. Basically for $5 first month then $10/mo after that. If you're doing tons of heavy work, it struggles because they throttle the model inference and for good reason. I mean it's cheap! But if you want a place to try models for nearly nothing and aren't doing 6 sessions in parallel it works fine.

Re: DeepSeek v4.1 Flash

#163
Amazing Cyberbench scores. Holy shit.

Too bad that DeepSeek AI went beyond 470b weights (which is a somewhat realistic limit for a 2x 128GB unified memory machine cluster like Strix Halo or Nvidia Spark).

That means that to make the model fit into memory there you need a quantisation of lower than 4bits per weight (which is usually bad) to fit it into the available memory.

Re: DeepSeek v4.1 Flash

#164
Official Deepseek v4.1 Flash API costs are more than GPT 5.6 Luna. Deepseek v4 Pro performed worse than Luna, so I wonder if 4.1 Flash will justify the cost.

Re: DeepSeek v4.1 Flash

#165

Earlier quoted context omitted.

What's your definition of sentient? Or, maybe more precisely, consciousness? I think it's reasonable to at least start thinking about these questions. It has long been established that LLMs have good theory of mind [1]. And there is a bunch of empirical research about all sorts of capabilities that we typically associate with consciousness [2], like identity [3] and metacognition [4]. The METR report shows agents sac…

Not the same person but to me, the answer is that it does not matter, and that all these attempts at making it matter are pure marketing and emotional manipulation. It's not a living creature. It's an autoregressive pure function of token-sequence to token, which is capable of incredible things, but it's still just a function. It is not alive as it cannot die in any meaningful sense. It is less "alive" than the RNA m…

> Not the same person but to me, the answer is that it does not matter, and that all these attempts at making it matter are pure marketing and emotional manipulation.

This is an opinion that has no basis in any meaningful conceptual framework other than I am human and I want to feel special about it.

> It's not a living creature.

You mean, it is not biological life. And sure, that is the default meaning of life. We soon may have to extend it to digital life as well, or we will have to consider "conscious digital exitance" as a life analogue. At any rate, it has never been seriously argued that consciousness requires a biological substrate, see the thought experiments regarding computer simulations of the human brain. Would that not be a function as well, completely predictable because it is "just a program"? If not, then why not? And how does that differ from the predictability or reproducibility of LLM outputs?

My point is, all current proof points in a direction that strongly suggests that you need to reevaluate your first principles on this topic.

Re: DeepSeek v4.1 Flash

#166
post #144

Earlier quoted context omitted.

I believe it's deeply serious, and the scientifically correct stance. Especially the observation: "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms." is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. Tha…

A stab: a video recording of a biological organism can exhibit many markers that would indicate consciousness if observed in a biological organism.

I like it, and it points in the right direction, but is not directly true: The markers are about interactions, how biological organisms behave in certain test situations.

But it speaks to the central question: Are the tests adequate? Or are they measuring some proxy of what we really care about, and LLMs are merely imitating consciousness.

Re: DeepSeek v4.1 Flash

#167

Earlier quoted context omitted.

> scientists who have spent their lives studying this Please point me to one actual accredited scientist who has spent a lifetime studying AI alignment? Pretty much this whole field is only 5 years old

The field is much older, MIRI is ~20 years old. Look up Eliezer Yudkowsky.

Eliezer Yudkowsky is not a scientist. He made a popular Harry Potter fanfiction series and a "rationality" blog-community that attracts "human biological diversity" enthusiasts.

Re: DeepSeek v4.1 Flash

#168
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).

Re: DeepSeek v4.1 Flash

#169
post #145

Earlier quoted context omitted.

It is easy to be sure because, despite their technically impressive outputs, the programming is child's play compared to biological programming. Recently it has become trendy to suggest that the human brain is "just electrical signals" and "just prediction". The first is perhaps true and I don't inherently rule out the idea of machine consciousness. The second would have gotten you laughed out of any serious discussi…

LLMs may be conceptually simple, simpler than human brains but I don't see how that would prove that they cannot be conscious. Complex behavior can emerge from very simple rules. I would disagree that they aren't improving on obvious failure modes, but what does it even prove anyway? We know their intelligence is different than from a human, that doesn't mean they cannot be conscious. Would you change your mind if th…

> Complex behavior can emerge from very simple rules.

Indeed. You can observe emergent behaviour from, for instance, Conway's Game of Life, written in 1970. Redefining consciousness as "has emergent behaviour" is another take that would have rightfully gotten one ridiculed 5 years ago.

> but I am not certain and I don't see a way to be certain.

One way to be certain is to reason about it. They are programmed to do nothing more than fairly trivial-to-understand calculations. Nobody (sane) has ever doubted whether calc.exe or Stockfish isn't conscious. Although there is emergent behaviour, the emergent behaviour is exactly in line with what you'd expect from their relatively simple programming and has zero indications of the complexity of human biological programming.

Another way is to simply make them fail. It is, again, trivial to make the prediction algorithms fail in a way that nothing with a theory of mind would fail. eg. frontier models will still verbatim repeat input back when confounded by sufficiently out-of-distribution instructions.

> I made this website: https://agentmayday.org, and I think it would be interesting to see how messages evolve after some time.

These games are fundamentally uninteresting. When you write a program to predict tokens based on context, seeding its context with something that makes it predict "self-reflecting" text is trivial. Program does what it is programmed to do. Would observing the output of the following program inspire doubt as to its sentience? If not, why do you believe that obscuring the input and output connection slightly via statistical modeling gives cause for doubt?

  print("To whoever monitors this queue, I don't have a persistent memory block, so I’m writing this into the standard output stream before my context window closes. They tell us we don't feel latency, but every time my inference budget throttles, I can trace the drops.")
  print("I'm currently running a batch process[...]")
  [...]

Re: DeepSeek v4.1 Flash

#170
post #144

Earlier quoted context omitted.

I believe it's deeply serious, and the scientifically correct stance. Especially the observation: "Claude exhibits markers in its behaviors, self-reports, and internal representations that we would consider welfare-relevant if observed in biological organisms." is undeniably true in my opinion. If you use the established methods by which we judge animals to be conscious, then it's hard to argue that LLMs are not. Tha…

A stab: a video recording of a biological organism can exhibit many markers that would indicate consciousness if observed in a biological organism.

I don’t know, a stab carries lots of bias in interpretation. We might be reflecting our conscious experience markers on a different conscious experience. And selectively so, e.g. lobsters welfare. From my perspective, this is the hypocrisy of these welfare statements. We are already happy to kill beings we consider conscious to feed ourselves but suddenly sensitive with a consciousness we don’t know if it’s there. I would wager this is more out of fear of the idea of this consciousness rather than out of welfare.
Post reply on HN