Live data from Hacker News

DeepSeek-v3.1-Terminus

api-docs.deepseek.com

1–10 of 30 posts

Re: DeepSeek-v3.1-Terminus

#2
> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars.

It's good that they made this improvement. But is there any advantages at this point using DeepSeek over Qwen?

Re: DeepSeek-v3.1-Terminus

#4
post #2

> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars. It's good that they made this improvement. But is there any advantages at this point using DeepSeek over Qwen?

I wish there was some easy resource to keep up with the latest models. The best I have come up with so far is asking one model to research the others. Realistically I want to know latest versions, best use case, performance (in terms of speed) relative to some baseline, and hardware requirements to run it.

Re: DeepSeek-v3.1-Terminus

#5
post #2

> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars. It's good that they made this improvement. But is there any advantages at this point using DeepSeek over Qwen?

MIT license that lets you run it on your own hardware and make money off of it.

Re: DeepSeek-v3.1-Terminus

#7
post #3

I see no article in the link, just "news250922" header with some layout

It’s up again, check it.

Twitter/X post link: https://twitter.com/deepseek_ai/status/1970117808035074215

Also Hugging Face model link: https://huggingface.co/deepseek-ai/DeepSeek-V3.1-Terminus

Re: DeepSeek-v3.1-Terminus

#8
post #2

> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars. It's good that they made this improvement. But is there any advantages at this point using DeepSeek over Qwen?

I wish there was some easy resource to keep up with the latest models. The best I have come up with so far is asking one model to research the others. Realistically I want to know latest versions, best use case, performance (in terms of speed) relative to some baseline, and hardware requirements to run it.

> asking one model to research the others.

that's basically choosing are random with extra steps!

Re: DeepSeek-v3.1-Terminus

#10
post #8

Earlier quoted context omitted.

I wish there was some easy resource to keep up with the latest models. The best I have come up with so far is asking one model to research the others. Realistically I want to know latest versions, best use case, performance (in terms of speed) relative to some baseline, and hardware requirements to run it.

> asking one model to research the others. that's basically choosing are random with extra steps!

Research not spit out the answer based on weights. Just ask Gemini/Claude to do deep research on /r/LocalLLama and HN posts.
Post reply on HN