Live data from Hacker News

Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

phoronix.com

661–670 of 722 posts

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#661

Earlier quoted context omitted.

How do architectural bottlenecks due to modified Von Neumann architectures' debuggable instruction pipelines limit computational performance when scaling to larger amounts of off-chip RAM? Tomasulo's algorithm also centralizes on a common data bus (the CPU-RAM data bus) which is a bottleneck that must scale with the amount of RAM. Can in-RAM computation solve for error correction without redundant computation and con…

For whatever reason Hynix hasn't turned their PIM into a usable product. LPDDR based PIM is insanely effective for inference. I can't stress this enough. An NPU+LPDDR6 PIM would kill GPUs for inference.

How many TOPS/W and TFLOPS/W? (T [Float] Operations Per Second per Watt (hour *?))

/? TOPS/W and FLOPS/W: https://www.google.com/search?q=TOPS%2FW+and+FLOPS%2FW :

- "Why TOPS/W is a bad unit to benchmark next-gen AI chips" (2020) https://medium.com/@aron.kirschen/why-tops-w-is-a-bad-unit-t... :

> The simplest method therefore would be to use TOPS/W for digital approaches in future, but to use TOPS-B/W for analogue in-memory computing approaches!

> TOPS-8/W

> [ IEEE should spec this benchmark metric ]

- "A guide to AI TOPS and NPU performance metrics" (2024) https://www.qualcomm.com/news/onq/2024/04/a-guide-to-ai-tops... :

> TOPS = 2 × MAC unit count × Frequency / 1 trillion

- "Looking Beyond TOPS/W: How To Really Compare NPU Performance" (2023) https://semiengineering.com/looking-beyond-tops-w-how-to-rea... :

> TOPS = MACs * Frequency * 2

> [ { Frequency, NNs employed, Precision, Sparsity and Pruning, Process node, Memory and Power Consumption, utilization} for more representative variants of TOPS/W metric ]

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#662

Earlier quoted context omitted.

Why do you need two DP cables? Is there not enough bandwidth in a single one? I use a 4k@60 display, which is the maximum my cheap Anker USB-C Hub can manage.

I'm not sure, but there's an in-depth exploration of the monitor here: https://tftcentral.co.uk/reviews/acer_nitro_xv273k.htm Reddit also seems to have some people who have managed to get 144 with FreeSync, but I've only managed 120. Funnily enough while I was typing this Netflix caused both my monitors to blackscreen (some sort of NVIDIA reset I think) and then come back. It's not totally stable!

This is likely a cable issue. Certain cable types can't handle 4k. I had to switch from DisplayPort to HDMI with a properly rated cable to get past this in the past.

It works up until too many pixels change, basically.

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#664

Earlier quoted context omitted.

Why do you need two DP cables? Is there not enough bandwidth in a single one? I use a 4k@60 display, which is the maximum my cheap Anker USB-C Hub can manage.

If you're on an Nvidia 4000 series DisplayPort is limited to 1.4, ~26 Gigabit/s. https://linustechtips.com/topic/729232-guide-to-display-cabl... is a calculator for bandwidth 4K@144 HDR is ~40 Gigabit/s. You can do better with compression, but I find Nvidia cards have an issue with compression enabled.

Interesting. I've been running my 4K monitor at 240Hz with HDR enabled for months and haven't had any issues with Display Stream Compression on my 4080.

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#665

Earlier quoted context omitted.

Curious as I’m of the same mind - what’s your local AI setup? I’m looking to implement a local system that would ideally accommodate voice chat. I know the answer depends on my use case - mostly searching and analysis of personal documents - but would love to hear how you’ve implemented.

If you are just starting up, you can try out 'open-webui' as inspiration. After that you can just use llama.cpp to build out your own things. Hardware side, I just have a beefy server that acts as a router (mellanox card to provider fiber optic and local fiber network), firewall, wifi access point, zigbee coordinator, host to various services, camera video feed ingestion and processing, and so on...

define brefy server

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#666

12GB memory -.- I feel like _anyone_ who can pump out GPU's with 24GB+ of memory that are capable to use for py-stuff would benefit greatly. Even if it's not as performant as the NVIDIA options - just to be able to get the models to run, at whatever speed. They would fly off the shelves.

Would it though? How many people are running inference at home? Outside of enthusiasts I don't know anyone. Even companies don't self-host models and prefer to use APIs. Not that I wouldn't like a consumer GPU with tons of VRAM, but I think that the market for it is quite small for companies to invest building it. If you bother to look at Steam's hardware stats you'll notice that only a small percentage is using high…

    Would it though? How many people are running inference at home? 
I don't know how to quantify it, but it certainly seems like a lot of people are buying consumer nVidia GPUs for compute and the relatively paltry amounts of RAM on those cards seems to be the number one complaint.

So I would say that Intel's potential market is "everybody who is currently buying nVidia GPUs for compute."

nVidia's stingy consumer RAM choices also seem to be a fairly transparent ploy to create a protective moat around their insanely-high-profit-margin datacenter GPUs. So that just seems like kind of an obvious thing for Intel or AMD to consider tackling.

(Although, it has to be said, a lot of commenters have pointed out that it's not as easy as just slapping more RAM chips onto the GPU boards; you need wider data busses as well etc.)

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#667

Earlier quoted context omitted.

This is the weird part, I saw the same comments in other threads. People keep saying how everyone yearns for local LLMs… but other than hardcore enthusiasts it just sounds like a bad investment? Like it’s a smaller market than gaming GPUs. And by the time anyone runs them locally, you’ll have bigger/better models and GPUs coming out, so you won’t even be able to make use of them. Maybe the whole “indoctrinate users t…

The future of local LLMs is not people running it on their PCs. It's going to be a HomePod/AppleTV/Echo/Google Home -style box you set up in a corner and forget about it. Then your devices in the ecosystem can offload some LLM tasks to that local system for inference, without having to do everything on-device.

This makes sense in some ways technologically, but just having a "centralized compute box" seems like a lot more complexity than many/most would want in their homes.

I mean, everything could have been already working that way for a lot of years right? One big shared compute box in your house and everything else is a dumb screen? But few people roll that way, even nerds, so I don't see that becoming a thing for offloaded AI compute.

I also think that the future of consumer AI is going to be models trained/refined on your own data and habits, not just a box in your basement running stock ollama models. So I have some latency/bandwidth/storage/privacy questions when it comes to wirelessly and transparently offloading it to a magic AI box that sits next to my wireless router or w/e, versus running those same tasks on-device. To say nothing of consumer appetite for AI stuff that only works (or only works best) when you're on your home network.

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#668

Earlier quoted context omitted.

How do architectural bottlenecks due to modified Von Neumann architectures' debuggable instruction pipelines limit computational performance when scaling to larger amounts of off-chip RAM? Tomasulo's algorithm also centralizes on a common data bus (the CPU-RAM data bus) which is a bottleneck that must scale with the amount of RAM. Can in-RAM computation solve for error correction without redundant computation and con…

For whatever reason Hynix hasn't turned their PIM into a usable product. LPDDR based PIM is insanely effective for inference. I can't stress this enough. An NPU+LPDDR6 PIM would kill GPUs for inference.

Is this fast enough for DDR or SRAM RAM? "Breakthrough in avalanche-based amorphization reduces data storage energy 1e-9" (2024) https://news.ycombinator.com/item?id=42318944

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#669
post #579

Earlier quoted context omitted.

To get 128GB of RAM on a GPU you'd need at least a 1024 bit bus. GDDR6x is 16Gbit 32 pins, so you'd need 64 GDDR6x chips, which good luck even trying to fit that around the GPU die since traces need to be the same length, and you want to keep them as short as possible. There's also a good chance you can't run a clamshell setup so you'd have to double the bus width to 2048 because 32 GDDR6x chips would kick off way to…

You do not need a 1024-bit bus to put 128GB of some DDR variant on a GPU. You could do a 512-bit bus with dual rank memory. The 3090 had a 384-bit bus with dual rank memory and going to 512-bit from that is not much of a leap. This assumes you use 32Gbit chips, which will likely be available in the near future. Interestingly, the GDDR7 specification allows for 64Gbit chips: > the GDDR7 standard officially adds suppor…

Yeah, the idea that you're limited by bus width is kind of silly. If you're using ordinary DDR5 then consider that desktops can handle 192GB of memory with a 128-bit memory bus, implying that you get 576GB with a 384-bit bus and 768GB at 512-bit. That's before you even consider using registered memory, which is "more expensive" but not that much more expensive.

And if you want to have some real fun, cause "registered GDDR" to be a thing.

Re: Intel announces Arc B-series "Battlemage" discrete graphics with Linux support

#670
post #523

Earlier quoted context omitted.

I still don't understand why graphics cards haven't evolved to include sodimm slots so that the vram can be upgraded by the end user. At this point memory requirements vary so much from gamer to scientist so it would make more sense to offer compute packages with user-supplied memory. tl;dr GPU's need to transition from being add-in cards to being a sibling motherboard. A sisterboard? Not a daughter board.

One of the reasons GPUs can have multiples of CPU bandwidth is they avoid the difficulties of pluggable dimms - direct soldered can have much higher frequencies at lower power. It's one of the reasons why ARM Macbooks get great performance/watt, memory being even "closer" than mainboard soldered RAM so getting more of those benefits, though naturally less flexibility.

This makes sense. Do we need to begin exploring different socket configurations? IE, using something that more closely resembles a CPU socket versus a traditional RAM slot.
Post reply on HN