Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

491–500 of 526 posts

Re: DeepSeek v4.1 Flash

#491

Earlier quoted context omitted.

This is adapted from Microsoft research's YOCO. It was known for a while(2024!). Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM. Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".

why didn't Microsoft scale its own invention?

There are thousands of such techniques across different parts of the system. In ML, there are way too many ideas, and lots of people knowingly and unknowingly restate the same ideas. It's a new field, so even common language is not there. For an extreme example, so many improvements are restatements of 1960 signal processing techniques - obviously very few ML people have done DSP beyond the undergrad course. It is also highly empirical and many parts of deep learning (heck, even non-NN ML) are not understood yet.

Thus, the reality is that most of these ideas become polished only when its actually deployed and it has to work outside of a PoC. Since LLMs are a high capex product, only very few people actually make non-PoCs. Deepseek is in the business of low cost, fast inference. So they are the ones actually polishing these efficiency-ish ideas and combining many of them (this one, then engram which is based on multiple previous ideas including google brain's ngrammer) to make a coherent system. Openai and anthropic's systems will also involve a polished combination of multiple ideas for each of their systems - Luna is likely a combination of a few efficiency-ish ideas. Shame they won't publish though.

As for microsoft, they don't really sell models, they sell azure. So there is no reason for them to do the high capex scale out of these types of bags of techniques. In a sense, it did benefit them, others developed the model and now many US customers can serve DS4.1 Flash on Azure datacenters.

If it is not clear, I am not understating anything. Combining these rough ideas and making them work actually involves real novel ideas on top and is what is much more difficult than the academic results that were built upon. This also does not mean that the academic results are useless, they are what give us useful priors at all in what is a highly empirical field.

Re: DeepSeek v4.1 Flash

#492
post #463

Earlier quoted context omitted.

> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer. "You should be doing it just in case God ends up being real."

On the other hand, we have recently taught rocks to think about software engineering. And while not perfect, they're surprisingly good at it. Once you start building things that are even a little bit like minds, I suspect that it's worthwhile to consider that the future might end up looking a bit like science fiction. The alternative is to insist that Nothing Ever Happens, and the future won't get too weird. Which is…

God being real is also a possibility, so maybe we really should start praying. After all, he was allegedly making bushes and stones talk thousands of years before we did anything with thinking rocks.

Re: DeepSeek v4.1 Flash

#493

Earlier quoted context omitted.

I generally agree about the problem with anthropomorphizing. But I don't think Anthropic are doing that. They explicitly write "in biological entities this would be considered a sign of consciousness, but we don't know how to interpret it here". However, I disagree with your point that "it's an autoregressive function, thus it doesn't matter". Let me explain why: Assume I do a complete neurological scan of a brain. I…

The problem with this line of thinking is that modern computers are nothing like the brain. LLMs don't stand on their own, they have to be run on these modern computers, but doing so does not change the physical properties of the computer. The simulation you propose of the brain is likely impossible due to quantum mechanics making it impossible to fully simulate: https://en.wikipedia.org/wiki/Quantum_mind Perhaps we'…

At first I was inclined to agree with you, but then I realized that the brain requires this whole complicated contraption (the body) to run and, really, do anything at all. And while I'm not familiar with the notion of 'quantum mind' I do think that biological processes aren't deterministic (at a cellular level).

And I think this does mirror the situation with LLMs -- you need this whole computer contraption and GPU, also running on electricity, to support the LLM's "thought" processes. And that if we model the brain's neurology sufficiently (which it seems we've done) we can achieve results that appear to be like thinking, even if it is an emergent behavior from "relatively" simple math.

Which actually makes me wonder the opposite -- are we, as humans, not much better than these LLMs? Suppose the body is just that super complicated computer contraption, honed by thousands/millions of years of evolution to achieve some semblance of homeostasis? If you reject the idea that we have a soul, we start to look very similar to the machines we build. "You are a brain inside a skull cockpit, piloting a bone mech covered in meat armor and skin" feels more and more relevant. That I'm just a meat circuit running brain chips and once you pull the plug on the source of electricity it all just... stops

Re: DeepSeek v4.1 Flash

#494

Earlier quoted context omitted.

words like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"

The AI didn't reply anything. It doesn't know anything. That was a high probability token sequence based on the contents of the context window up to that point. It might well be correct, because the process for generating that next token distribution includes billions of parameters trained on, among other things, the entire body of LLM and transformer literature until the training data cut off. But that doesn't mean…

I don’t know, so is your brain?

And in the end it’s all elementary particles and four fundamental physical forces that even unify to one at high energy.

That sort of reductionism is kind of useless. “It’s cloudy outside, it makes me sad.” - “Oh bollocks, it’s just non-qualitative changes in wavelength and intensity.”

Re: DeepSeek v4.1 Flash

#495

Earlier quoted context omitted.

Yes, this is well-documented and publicly advertised. In Azure Foundry, the feature to modify (or completely remove) safety guardrails and content filtering is called "Limited Access" [0], and one must submit a form to request permission to use this feature. This is one of the more straightforward paths to get access to unrestricted frontier models, but it's far from the only way. [0] - https://learn.microsoft.com/en…

This looks like it removes additional guardrails put on by Microsoft, not native guardrails from OpenAi / Anthropic?

No. It's not limited to 3rd-party guardrails. Given how restrictive the native public-facing OpenAI guardrails are, this feature wouldn't be worth very much if it just slacked back off to the level of "regular" filter paranoia offered by the native models, would it?

This is needed if you're going to be dealing with things like psychologists doing self-harm research or red-teaming or sensitive sexual content -- if you're working with any of that sort of stuff in a professional context and want to leverage OpenAI models on Azure, then that's the form that you fill out to get access to unfiltered models.

Note that I am not aware of this feature being offered for Anthropic models -- I've only seen it offered for OpenAI models (note that the documentation I linked is specifically in the "Azure OpenAI" category).

Re: DeepSeek v4.1 Flash

#496
post #207

Earlier quoted context omitted.

> welfare We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer. "You should be doing it just in case God ends up being real."

Sure if you think that AI spiraling out of control is equally as likely as a magical fairy in the cosmos.

Re: DeepSeek v4.1 Flash

#497

Earlier quoted context omitted.

I always talk to models using grugspeak, like 'where getcontext used' I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this

That’s just how Chinese words work

Depending in what you are doing, you can unlock a lot more knowledge by translating your query to Chinese and then asking that. :)

Re: DeepSeek v4.1 Flash

#498

Earlier quoted context omitted.

words like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"

The AI didn't reply anything. It doesn't know anything. That was a high probability token sequence based on the contents of the context window up to that point. It might well be correct, because the process for generating that next token distribution includes billions of parameters trained on, among other things, the entire body of LLM and transformer literature until the training data cut off. But that doesn't mean…

Here's my GUT. The multiple branches of the Multiverse (MWI) are in competition to become the best branch. So they steal probabilities from each other since the probabilities have to sum upto 1. This competition could be considered as Painful.

Re: DeepSeek v4.1 Flash

#499
post #247
post #144

Earlier quoted context omitted.

A stab: a video recording of a biological organism can exhibit many markers that would indicate consciousness if observed in a biological organism.

A video is a fixed representation. What if we can interact with this video, and it reacts in the same ways the source organism does? Then we put it in new situations that weren't in the source video, and it interacts in a similar way to the original organism in these situations, too. What do we make of reactions of pain or joy? Where's the line between simulation and enaction? This is closer to the reality of these m…

The example can continue with a depiction of an organism in a video game
Post reply on HN