Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

31–40 of 426 posts

Re: DeepSeek v4.1 Flash

#32
post #18

If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.

Yes, this is one of the few issues with Deepseek; their chat pages and the app all respond in Chinese. However, i think i have only had it happen once when using the API, and im using it for hours each day for the last... couple of months?

Re: DeepSeek v4.1 Flash

#33
post #18

If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.

Same issue on desktop. Would be nice be able to set a prefix or postfix for every prompt.

Re: DeepSeek v4.1 Flash

#34
post #27

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

Bias…

Re: DeepSeek v4.1 Flash

#35
post #18

If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.

I just started learning Chinese instead, like they want us to

seriously

Re: DeepSeek v4.1 Flash

#36

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

Yes, a model's technical report should first and foremost include technical details.

Re: DeepSeek v4.1 Flash

#37

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

Wow there really is a model welfare section in there...

Re: DeepSeek v4.1 Flash

#38
post #27

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

Because safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivalent of a toddler.

Re: DeepSeek v4.1 Flash

#39

Initial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristicti…

> My favourite benchmark for this is to ask it to download a rom for an old game

Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked.

And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.

Re: DeepSeek v4.1 Flash

#40
post #18

If only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.

I occassionally get Chinese characters interlaced with English in Google AI Mode, too.
Post reply on HN