Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

541–543 of 543 posts

Re: DeepSeek v4.1 Flash

#541
post #263
post #161

I've run some evals on my puzzle game https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.

Since it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.

Thanks for the help. I ran it on high and it did pretty well and got a lot of one-shots in. The reasoning makes a much bigger difference than some other models.

Re: DeepSeek v4.1 Flash

#542

Earlier quoted context omitted.

> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer. "You should be doing it just in case God ends up being real."

Sure if you think that AI spiraling out of control is equally as likely as a magical fairy in the cosmos.

I do, make of that what you will.

Re: DeepSeek v4.1 Flash

#543

Earlier quoted context omitted.

God being real is also a possibility, so maybe we really should start praying. After all, he was allegedly making bushes and stones talk thousands of years before we did anything with thinking rocks.

Its only a possibility if you reject modern science.

What science has proven, beyond the shadow of a doubt, that nothing comes after death? I'm sure most of the human race would be very interested to read the white paper.
Post reply on HN