Live data from Hacker News

Intel Arc Pro B70 Review

pugetsystems.com

101–110 of 131 posts

Re: Intel Arc Pro B70 Review

#101
post #76

Intel Arc B70 when released, can only produce 1/3 of the token of RTX PRO 4500. Well, it also cost 1/3 of RTX PRO 4500. It lacked software support the for the primary target application, running LLM. The officially supported vllm fork is 6 version behind mainline. It did not run the latest hot new open models on huggingface. Parallel two of B70 reduce token rate, not improve it. So, the software behind B70 is basical…

What you say is not consistent with TFA.

The parent article shows that B70 is faster than RTX 4000.

RTX 4500 is faster than RTX 4000, but it cannot be more than 3 times faster, not even more than 2 times faster.

The parent article is consistent with RTX 4500 being faster than B70 for ML inference, but by a much smaller ratio, e.g. less than 50% faster.

If you know otherwise, please point to the source.

If you have run a benchmark yourself, please describe the exact conditions.

In the benchmarks shown at Phoronix for llama.cpp, the relative performance was extremely variable for different LLMs, i.e. for some LLMs a B70 was faster than RTX 4000, but for others it was significantly slower.

Your 3x performance ratio may be true for a particular LLM with a certain quantization, but false for other LLMs or other quantizations.

This performance variability may be caused by immature software for B70. For instance instead of using matrix operations (XMX engines), non-optimized software might use traditional vector operations, which are slower.

It is also possible that for optimum performance with a certain LLM one may need to choose a different quantization for B70 than for NVIDIA, because for sub-16-bit number formats Intel supports only integer numbers.

Re: Intel Arc Pro B70 Review

#102

Earlier quoted context omitted.

I still not see the point running these models. I say they produce plausible garbage, nowhere near quality of frontier models (when they work). Why can't Intel look beyond this nonsense state of affair and build something with 1TB of RAM or more? What I am trying to say, I am yet to see anything competitive in the market. Cards very much stalled in sub 100GB region and best corporations can do is throw something to r…

What's wrong with Grace Hopper if you want to throw buckets of local memory at a problem?

Some people, including myself, loathe Nvidia with the fiery burning passion of a thousand suns, and will put up with whatever nonsense is necessary to run without them.

Re: Intel Arc Pro B70 Review

#103
post #91
post #78

Earlier quoted context omitted.

There are nonlinearities to exploit in that calculus. Given enough VRAM to host a larger model that you're targeting, just the size can push you past the usability threshold at a much better price.

Problem is the more B70 you have, the slower the inference it gets(due to terrible software atm). A single B70 is almost barely faster than CPU inference. If you have 4 B70, you might as well run interference on CPU and be faster with cheaper DDR5 instead of GDDR6.

For what you say to be useful, please specify what sowftware you have used with B70, including its version.

Hardware-wise a B70 should be significantly faster than any of the available CPUs at ML inference. If it was not so in your tests, that must really be a software problem, so you must identify the software, for others to know what does not work.

Re: Intel Arc Pro B70 Review

#104

There's a tradeoff between dense models and MoEs on memory usage vs. compute for the same quality. For example, Qwen3.5 27B and Qwen3.5 122B A10B have similar average performance across benchmarks. The 122B is much faster to run than the 27B (generates more tokens at the same compute). The 27B, on the other hand, uses ~4x less VRAM at low context lengths (less difference at high context lengths). Right now, different…

LLMs are memory bandwidth bound not compute bound.

LLMs are bound by both and depends on the hardware which factor is higher.

Re: Intel Arc Pro B70 Review

#105
post #17
post #9

Earlier quoted context omitted.

The drivers often need per game optimisations these will be missing but I doubt Intel would nerf them, just rely on you not paying a lot for RAM the game won't use.

I actually meant it in a different way. I would get it for local AI stuff, but being able to game on it would be a huge plus, otherwise I would need two different machines.

It'll work just fine for gaming. It's what the B770 would have been if it had 32GB RAM and ever got released.

Re: Intel Arc Pro B70 Review

#106

Earlier quoted context omitted.

I still not see the point running these models. I say they produce plausible garbage, nowhere near quality of frontier models (when they work). Why can't Intel look beyond this nonsense state of affair and build something with 1TB of RAM or more? What I am trying to say, I am yet to see anything competitive in the market. Cards very much stalled in sub 100GB region and best corporations can do is throw something to r…

What's wrong with Grace Hopper if you want to throw buckets of local memory at a problem?

Most consumer platforms only allow up to 128/256GB of RAM. If you want more you likely need a data centre platform. This is again a mismatch between what companies think consumers are at and the reality.

I think e.g. AMD missed the boat with 9950x3d2 by limiting memory controller. If it was possible to hook it with 1TB of consumer DDR5 RAM, that would be something to write home about.

Re: Intel Arc Pro B70 Review

#107
Can we not have a PCIe card that's ASIC (and isn't GPU) with even DDR 4 or DDR 5 memory (Let's say 128 GB) onboard and being able to shove four of them on a consumer grade motherboard and then being utilized in parallel?

Noob question.

Re: Intel Arc Pro B70 Review

#108
Something that is also cool with these cards is proper SR-IOV without hassle. Arc pro cards make for nice graphical acceleration devices for vms. I know ai gets all the hype but I also appreciate being able to accelerate multiple workstations with a single gpu and still get decent frametimes.

Re: Intel Arc Pro B70 Review

#109
post #27

I was looking into this for LLMs but it's clearly a graphics-processing focused card. The memory bandwidth is too low for that much RAM to be useful in an LLM context. The 5090 I have has the same amount of RAM but far more bandwidth and that makes it much more useful.

> it's clearly a graphics-processing focused card.

Yes, that's what the G in GPU stands for. It's great to see that there are still manufacturers that understand this.

Re: Intel Arc Pro B70 Review

#110
post #100
post #17

Earlier quoted context omitted.

I actually meant it in a different way. I would get it for local AI stuff, but being able to game on it would be a huge plus, otherwise I would need two different machines.

Much as I want diversity; a 3090 would be a billion times better for games and can probably hold its own for a broader AI workload. Anything other then running highly quantised models that don't fit in 24GB with realativly small contexts.

A 3090 is what I have now.

But I hope to somehow have 48Gb or 64GB VRAM in a GPU that's also gaming-ready.

I was looking for maybe getting a mac studio for this reason, but I don't think a mac is really good for for gaming.

Post reply on HN