Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

471–480 of 514 posts

Re: DeepSeek v4.1 Flash

#471

Earlier quoted context omitted.

The field was purely theoretical 20 years ago, and Yudkowsky is pretty much the dictionary definition of "not accredited"

some use "not accredited" as a pejorative term. Lets not forget that the 'Fermat's Last Theorem' which has been pretty visible for the non-math crowd of late due to the recent AI frenzy about a purported proof was but one small contribution to the world's math lexicon by someone with a bachelors degree in civil law, that George Green was a baker and millwright, Boole was the son of a poor shoemaker in England with no…

> some use "not accredited" as a pejorative term.

I am absolutely using "not accredited" in a pejorative sense here. That he publishes papers coauthored by a couple of philosophy professors at Oxford (all of whom have made a ton of money from the Silicon Valley AI Alignment and Effective Altruism crowds) does not make him a scientist.

I will also note that the Oxford Philosophy department finally shitcanned the whole Future of Humanity Institute a couple of years back.

Re: DeepSeek v4.1 Flash

#472
post #388

Earlier quoted context omitted.

No, that is not at all something we can just state as a fact. Whether the brain is deterministic is an open question that just inherits the good old, probably unsolvable determinism debate. The LLM pseudo-randomness from above is engineered by us humans and fully understood, much like an algorithm playing a video frame sequence. You could theoretically record a full register of all states of an LLM setup with all the…

Nonsense. There was one proposal for relevant quantum effects in brain dynamics, and that turned out to be not relevant. Even if they were, you could substitute all quantum randomness with pseudo randomness and obtain an absolutely indistinguishable object. But even if this were a debate, its absolutely absurd to claim that the question of determinism in the brain has any bearing on our moral standing. If we discover…

You seem to have conceded the "is it deterministic" argument only to sidestep by declaring determinism irrelevant. Your original claim was that LLMs are deterministic "in the same sense" as brains.

We can write down an LLMs full register, and that register/book contains the whole output universe of the text generator. That book does not act, it is morally neutral. That the brain has such a register at all is just restating the determinism axiom, which you treat as fact.

A text is not conscious, and we can not wish it into consciousness, no matter how many human-like patterns we find in the book / the generated text. It has not been shown that running the text adds anything over the text written out. Researchers are super motivated to find machine consciousness but cannot find it, while a company months from its IPO keeps pitching shadows of consciousness all day. It really is a PR strategy.

Re: DeepSeek v4.1 Flash

#473
post #274

Earlier quoted context omitted.

No language models are programmed, they are "grown" or evolved from data. There's no print statements or human entered logic involved in the raw model expression at all. The only thing that humans have programmed is efficient parallel dot product pipelines that "animate" (for lack of a better word) the models. Everything these models do is emergent from their backpropgation guided evolution. This even includes in con…

You have completely misunderstood what I was saying so badly I can't even formulate a response other than to suggest you read my reply again. I was not suggesting that LLMs are programmed with print statements, for fuck's sake.

If you say so.

> When you write a program to predict tokens based on context, seeding its context with something that makes it predict "self-reflecting" text is trivial. Program does what it is programmed to do. Would observing the output of the following program inspire doubt as to its sentience?

Then you follow it up with print statements as if that is a good analogy.

As I said, they are not programmed, so your question above is not relevant to your argument.

You say they're programs that are stochastically jiggled, but that's simply not accurate either. All LLM abilities are emergent, even when the training corpus is well defined.

I didn't think you literally thought they were made of print statements, but you are implying they're software that's been "fuzzed". Hopefully you don't literally that either and you're just using it as a bad analogy.

You could have argued from the stance of neural networks being universal functions, which might at least be closer to the truth, but instead your example is print statements!

I get you're trying to say that something trained to say a thing doesn't mean it has arrived at the thing like a mind would, and perhaps that would have been closer for GPT 2.

These days though, we just have so much more awareness of what they're actually doing internally that it's bizarre to even compare them to stochastic parrots of the training corpus, if that is closer to what you're implying.

For example: https://www.anthropic.com/research/global-workspace

https://transformer-circuits.pub/2025/attribution-graphs/bio...

Re: DeepSeek v4.1 Flash

#477

Earlier quoted context omitted.

> Much more knowledgeable than you or I are Speak for yourself. I work for an LLM startup that was successfully bootstrapped and is now highly profitable with 8-digit revenue and zero outside investment. Unlike OpenAI and Anthropic, we do not rely on deceiving investors to dump a trillion dollars into a tar fire with the false promise of delivering the machine god that will unemploy all of humanity (at best). Taking…

> Speak for yourself. I work for an LLM startup And yet you still fail to demonstrate good understanding of the topic ¯\_(ツ)_/¯ > stating that those are the only people who can be trusted You are right, they are most definitely not the only people who can be trusted to have current and accurate information. But due to the unique constraints of these fast-moving events, they are certainly among those whose opinions ne…

Simple cellular automata demonstrate emergent behavior. Emergent behavior is nothing new in computer science and is not remotely unique to LLMs.

Re: DeepSeek v4.1 Flash

#478

Earlier quoted context omitted.

I always talk to models using grugspeak, like 'where getcontext used' I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this

words like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"

The AI didn't reply anything. It doesn't know anything. That was a high probability token sequence based on the contents of the context window up to that point. It might well be correct, because the process for generating that next token distribution includes billions of parameters trained on, among other things, the entire body of LLM and transformer literature until the training data cut off. But that doesn't mean that AI holds any particular opinion about anything. The reply is the opinion of the pretraining data and the subsequent rounds of RL not of a conscious artificial intelligence as such.

Re: DeepSeek v4.1 Flash

#479
post #274

Earlier quoted context omitted.

> Complex behavior can emerge from very simple rules. Indeed. You can observe emergent behaviour from, for instance, Conway's Game of Life, written in 1970. Redefining consciousness as "has emergent behaviour" is another take that would have rightfully gotten one ridiculed 5 years ago. > but I am not certain and I don't see a way to be certain. One way to be certain is to reason about it. They are programmed to do no…

No language models are programmed, they are "grown" or evolved from data. There's no print statements or human entered logic involved in the raw model expression at all. The only thing that humans have programmed is efficient parallel dot product pipelines that "animate" (for lack of a better word) the models. Everything these models do is emergent from their backpropgation guided evolution. This even includes in con…

They aren't grown/evolved from data, they are fit to the data. The fitting process can be fully deterministic although its fairly easy to screw things up such that it isn't deterministic, but that just a defect not some fundamental shift.

Re: DeepSeek v4.1 Flash

#480
post #287

I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost? Here's the same token usage priced at different rates: a real…

I ran the preview model around 2,126,605,070 tokens for $22.04 USD for the last couple of days. Kind of shocked. It did a decent job refactoring https://github.com/mmastrac/diffgemma to create a CUDA support backbone, it's struggling a bit to port metal kernels to CUDA unattended (it hasn't managed to get numbers to match over >1 layer). It successfully ported a root exploit to an older Android phone that GLM5.3Flash…

Which harness are you using?
Post reply on HN