Live data from Hacker News

DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

github.com

101–110 of 124 posts

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#101
With the arrival of APUs for AI everyone is going to lose interest in GPUs real fast.

Why buy an overpriced Nvidia 4090 when you can get an AMD Halo Strix or Apple M3 Studio APU with 512GB or 128GB of Ram?

Nvidia has kept prices high and performance low for as long as it can and finally competition is here.

Even Intel can make APUs with tons of RAM.

Nvidia hopefully is squirming.

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#103

I'm sure this question has been asked before, but why not launch a GPU with more but slower ram? That would fit bigger models while still affordable...

Did you miss the news about AMD Halo Strix?

More than twice as fast as Nvidia 4090 for AI.

Launched last week.

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#105
post #59

As other commenters have mentioned, the performance of this set up is probably not really great since there's not enough VRAM and lots of bits have to be moved between CPU and GPU RAM. That said, there are sub-256GB quants of DeepSeek-R1 out there (not the distilled versions). See https://unsloth.ai/blog/deepseekr1-dynamic I can't quantify the difference between these and the full FP8 versions of DSR1, but I've been…

Currently >8 token/s; there is a demo in this post: https://www.linkedin.com/posts/jasondai_run-671b-deepseek-r1...

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#107

Earlier quoted context omitted.

This article keeps getting posted but it runs a thinking model at 3-4 tokens/s. You might as well take a vacation if you ask it a question. It’s a gimmick and not a real solution.

If you value local compute and don't need massive speed, that's still twice as fast as most people can type.

Reasoning models spend a whole bunch of time reasoning before returning an answer. I was toying with QWQ 32B last night and ran into one question I gave it where it spent 18 minutes at 13tok/s in the phase before returning a final answer. I value local compute but reasoning models aren’t terribly feasible at this speed since you don’t really need to see the first 90% of their thinking output.

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#108

I'm sure this question has been asked before, but why not launch a GPU with more but slower ram? That would fit bigger models while still affordable...

Did you miss the news about AMD Halo Strix? More than twice as fast as Nvidia 4090 for AI. Launched last week.

I indeed was not aware, thanks

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#109

Earlier quoted context omitted.

What a teaser article! All this info for setting up the system, but no performance numbers.

That's because the OP is linking to the quickstart guide. There are benchmark numbers on the github's root page, but it does not appear to include the new deepseek yet: https://github.com/intel/ipex-llm/tree/main?tab=readme-ov-fi...

Am I missing something ? I see a lot of the small-scale models results but no results for DeepSeek-R1-671B-Q4_K_M on their github repos.

Re: DeepSeek-R1-671B-Q4_K_M with 1 or 2 Arc A770 on Xeon

#110

I'm sure this question has been asked before, but why not launch a GPU with more but slower ram? That would fit bigger models while still affordable...

Did you miss the news about AMD Halo Strix? More than twice as fast as Nvidia 4090 for AI. Launched last week.

> More than twice as fast as Nvidia 4090 for AI.

Not in memory bandwidth which is all that matter for LLM inference.

Post reply on HN