Live data from Hacker News

DeepSeek-v3.1

api-docs.deepseek.com

271–273 of 273 posts

Re: DeepSeek-v3.1

#271
post #255

Earlier quoted context omitted.

Currently no, but I'm running them! Some people on the aider discord are running some benchmarks!

@danielhanchen do you publish the benchmarks you run anywhere?

We had benchmarks for Llama 4 and Gemma 3 at https://docs.unsloth.ai/basics/unsloth-dynamic-2.0-ggufs - for others I normally refer to https://discord.com/channels/1131200896827654144/12822404236... which is the Aider Polygot Discord - they always benchmark our quants :)

Re: DeepSeek-v3.1

#272
post #29

Earlier quoted context omitted.

Like 4. Definitely single digit. The P40s are slow af

P40 has memory bandwidth of 346GB/s which means it should be able to do around 14+ t/s running a 24 GB model+context.

Not sure why I got downvoted - literally the first result (for me) says the best result[0] is 11t/s at Q3. Everything else is single digits, like 2-8t/s. Also considering that its not supported anymore[1] (It's Compute Capability is 6.1, not supported by cuda anymore) and it's power draw, I'd highly recommend anyone interested in ML stay far away from it - even if its all you can afford.

While the memory bandwidth is decent, you do actually need to do matmuls and other compute operations for LLMs, which again its pretty slow at

[0]: https://old.reddit.com/r/LocalLLaMA/comments/1dcdit2/p40_ben... [1]: https://developer.nvidia.com/cuda-gpus

Re: DeepSeek-v3.1

#273
post #50

Earlier quoted context omitted.

tbh companies like anthopic, openai, create custom agents for specific benchmarks

Aren't good benchmarks supposed to be secret?

if you're able to submit multiple times, then you can learn from the hidden set
Post reply on HN