I've run some evals on my puzzle game https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.
Since it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.
DeepSeek v4.1 Flash
541–544 of 544 posts
Re: DeepSeek v4.1 Flash
#542Earlier quoted context omitted.
> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance. This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer. "You should be doing it just in case God ends up being real."
Sure if you think that AI spiraling out of control is equally as likely as a magical fairy in the cosmos.
Re: DeepSeek v4.1 Flash
#543Earlier quoted context omitted.
God being real is also a possibility, so maybe we really should start praying. After all, he was allegedly making bushes and stones talk thousands of years before we did anything with thinking rocks.
Its only a possibility if you reject modern science.
Re: DeepSeek v4.1 Flash
#544Earlier quoted context omitted.
In what way does your brain change state while sleeping that a model does not change state via constant fine-tune updates?
It's possible that an LLM is conscious during training, but there are no "constant fine-tune updates" during inference.