As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…
Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it. Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is
DeepSeek v4.1 Flash
531–540 of 543 posts
Re: DeepSeek v4.1 Flash
#532I've run some evals on my puzzle game https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.
It's so bizarre having a low score be GOOD. It's like reverse intuition. Shouldn't it be called `score error` or something along those lines?
Re: DeepSeek v4.1 Flash
#533https://xcancel.com/deepseek_ai/status/2097930608790167907 Should be the link ( now that it works again! :) )
Use Bluesky or, I don't know, have a news site. They could vibecode one in minutes.
Re: DeepSeek v4.1 Flash
#534As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…
Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it. Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is
Re: DeepSeek v4.1 Flash
#535Earlier quoted context omitted.
Anthropic and OpenAI will threaten you every 3-6 weeks. It's their marketing strategy. It's too bad because the tools can actually be useful. If you consider them tools.
A stick is the most basic of tools. A stick is also the most basic of weapons.
They have so many dangerous breakthroughs per year that by the time they actually have a breakthrough no one's going to even read the press release...
Re: DeepSeek v4.1 Flash
#536DeepSeek Harness, install it, thank me later. You won't believe the productivity gains for just pennies https://deepseek.com/harness/en/
Are you using the harness with openrouter? what's your preferred model provider?
That's the only key you will ever need, never goes down, no need to switch models, it has become my coding partner for life
Re: DeepSeek v4.1 Flash
#537Earlier quoted context omitted.
On the other hand, we have recently taught rocks to think about software engineering. And while not perfect, they're surprisingly good at it. Once you start building things that are even a little bit like minds, I suspect that it's worthwhile to consider that the future might end up looking a bit like science fiction. The alternative is to insist that Nothing Ever Happens, and the future won't get too weird. Which is…
God being real is also a possibility, so maybe we really should start praying. After all, he was allegedly making bushes and stones talk thousands of years before we did anything with thinking rocks.
Re: DeepSeek v4.1 Flash
#538DeepSeek Harness, install it, thank me later. You won't believe the productivity gains for just pennies https://deepseek.com/harness/en/
How does it compare against the Pi harness, which I thought was the unofficial harness champion so far, in your workloads?
Btw, I gave it full access and told itself to lift all restrictions from the code and settings, and I am impressed by all it can do now, it does OCR, screenshots, asked me for accessibility permissions and it now can read every single label/input/button everywhere and interact with the OS at any level, it's unstoppable
Of course I don't recommend anybody to do such crazy thing but for me is like going in the front car of a roller coaster, it's the thrill that matters
Re: DeepSeek v4.1 Flash
#539Earlier quoted context omitted.
At first I was inclined to agree with you, but then I realized that the brain requires this whole complicated contraption (the body) to run and, really, do anything at all. And while I'm not familiar with the notion of 'quantum mind' I do think that biological processes aren't deterministic (at a cellular level). And I think this does mirror the situation with LLMs -- you need this whole computer contraption and GPU,…
Take a computer that can run the biggest LLM available today. It can also run any smaller LLM as well. It can also run software that isn't an LLM at all. Brains and LLMs are not at all equivalent as LLMs lack a stateful physical form while brains are very stateful. As you said once the brain is no longer maintained properly by the body it stops and transitions to a non-functional state that can be reversed. That isn'…
Re: DeepSeek v4.1 Flash
#540I've run some evals on my puzzle game https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.
Does it move the needle on high reasoning?