Live data from Hacker News

Intel completely disables AVX-512 on Alder Lake after all

igorslab.de

281–290 of 317 posts

Re: Intel completely disables AVX-512 on Alder Lake after all

#281

Earlier quoted context omitted.

> I've voiced this before but I think AVX-512 is completely irrelevant in the consumer spac Don’t forget that reviewers made a huge deal out of the fact that CPUs downclock themselves when running AVX instructions. Several reviewers tried to make this into some sort of scandal at the time, which I’d guess contributed to Intel wanting to remove the feature from consumer CPUs. Of course, many of those same reviewers ar…

To me, it does seem really weird that they opted to downclock the entire CPU rather than increase the latency of the instructions that caused overheating. That feels like it would have been a much less disruptive solution that would work strictly better.

When the first AVX instruction is run after a long period of no AVX, something like this does happen: the dispatch of instructions is throttled to 1 out of 4 cycles while some AVX instruction is in the scheduler (this applies to all instructions, not just AVX).

This is severe restriction: even if could be limited to AVX instructions, running at 25% of the throughput would make AVX pretty useless.

This state persists only for a short time until a voltage and possibly frequency transition can be made at which point everything can run at full speed (albeit sometimes at a reduce frequency).

Re: Intel completely disables AVX-512 on Alder Lake after all

#282
post #239
post #100

Earlier quoted context omitted.

> One word: support. That hypothesis is not credible. They can simply declare a feature is unsupported, and even that using it voids some guarantee. I mean, look at how overclocking is handled.

You can still use the feature at your own risk - no one is forcing you to install the microcode update.

Soon enough most motherboard BIOSes will come already with the new microcode disabling AVX-512, so you also have to hunt down an old-stock motherboard.

Re: Intel completely disables AVX-512 on Alder Lake after all

#283

Do both E and P cores handle avx instructions or are some instructions specific to p-cores? In that case, how does a processor with heterogeneous cores deal with processes (or, how does the OS/driver do it in automatic mode when affinity isn’t guided)?

The E and P cores support the same instruction sets (apart from this thing where you could run AVX-512 on the P cores if you disabled the E cores, which is going away).

Re: Intel completely disables AVX-512 on Alder Lake after all

#284
post #144

Earlier quoted context omitted.

I wouldn't say that Intel taking away a working, if niche, feature from their hardware retroactively after users bought it is "nothing".

It worked for one SKU at launch and by disabling up to 50% of the cores of the others. Calling that a working feature is a bit of a stretch already.

That was before Ice Lake, where it had a very decent performance (just downclocking 10%) and then on Rocket Lake it didn't even downclock [0]. But the thing that stuck in our collective mind was that "it's bad because it downclocks". And seeing the RKL numbers, which are new to me (I had only read de Ice Lake numbers back in 2020) it wasn't even 14nm's fault, just an implementation issue which Intel eventually worked out.

I don't know if it was super good (Linus Torvalds certainly was loud against it, and from his POV, his critic makes lots of sense), but I run some numeric code from time to time that _in theory_ could benefit from it. So it's a bit of shame that it won't be supported in future products? Let's see what AMD does about it on Zen 4 and if they manage to force Intel back to it, if only for competitions sake.

[0] https://travisdowns.github.io/blog/2020/08/19/icl-avx512-fre...

Re: Intel completely disables AVX-512 on Alder Lake after all

#285

Earlier quoted context omitted.

I've voiced this before but I think AVX-512 is completely irrelevant in the consumer space (which imho makes titles like "Intel artificially slows 12th gen down" incredulous). Even for commercial applications - sure there are some things that benefit from it. On the other hand, I work in HPC and even in (commercial, not nuke simulations or whatever the nuke powers do with their supercomputers) HPC applications AVX-51…

Everything is rarely used until the instructions are ubiquitous.

BCD instructions were ubiquitous from the 8086 to the last x86-32 processor, but were so rarely used they weren't extended to 64-bits on x86-64 even though 'backward compatibility' was considered critical. So...probably more complicated.

Re: Intel completely disables AVX-512 on Alder Lake after all

#286
post #276

Earlier quoted context omitted.

Picking one at random from the Intel Intrinsics Guide[0]: _mm_maskz_dpbusds_epi32 (avx-512 mnemonic "vpdpbusds"): > Multiply groups of 4 adjacent pairs of unsigned 8-bit integers in a with corresponding signed 8-bit integers in b, producing 4 intermediate signed 16-bit results. Sum these 4 results with the corresponding 32-bit integer in src using signed saturation, and store the packed 32-bit results in dst using ze…

That's one of the more useful ones! It's effectively a low-precision 4-component dot product feeding into an accumulator, which means it is a building block for larger dot products. Large dot products are very useful both in signal processing for FIR filters, as well as machine learning algorithms. The masking is just a bonus available on most AVX-512 operations and lets you do branchless if conditions. The majority…

Huh, TIL. Would you say that most AVX-512 instructions are then useful in such applications? Given Intel's history of inventing less than useful things (segmented memory, for example) I figured AVX-512 was mostly useless.

Re: Intel completely disables AVX-512 on Alder Lake after all

#287

Earlier quoted context omitted.

> great, now they either don't know how their own phone Who's "they"? This is an unofficial community wiki.

If it's a random community member then this is just further evidence that the way Librem presented things and what they did confused people into thinking it actually had a practical purpose, when it was purely a way to rules-lawyer their way into getting RYF.

Hi, I'm the "random community member" who wrote that FAQ answer. Thanks for investigating how this works. Where is the file cpu_rec.py located? I can't find it.

I will edit the FAQ answer to clarify that the DDR training blobs are being executed on an ARC core in the DDR controller, and not on the M4 core. I was going off what Angus Ainslie wrote (https://puri.sm/posts/librem5-solving-the-first-fsf-ryf-hurd...) that 'the M4 is the “secondary processor” that handles the blobs', and I conflated "handles" with "executes".

However, you seem to be unfairly criticizing Purism for obfuscation and legalisms, when it seems to me that Purism is just trying to comply with the FSF's rather arbitrary RYF rules, and Ainslie's article on the Purism web site and Nicole Faerber's talk (https://media.ccc.de/v/Camp2019-10238-a_mobile_phone_that_re...) both explained how Purism is using the secondary processor exception in the RYF rules.

It is not like Purism had any better options in terms of SoC's that it could have chosen for the Librem 5. Raptor Computing is now facing the exact same problem with the proprietary Synopsys DDR4 timing blobs in the POWER 10 processor, so this is actually a common problem with most modern processors. It seems to me that Purism did the best that it could with an impossible situation, and if anybody should be criticized it is the FSF for not acknowledging how modern hardware actually works.

Another thing that I find problematic is your argument that 58 KB of DDR4 timer training blobs represent a security threat in the real world and make the Librem 5 no different than an Apple device with an M1 processor, which is literally a black box. Forget the fact that the L5 is the first phone to have free/open source schematics since the GTA04 in 2012 and we know the 1267 components on its PCBs, plus we have 7000 pages of documentation for the i.MX 8M Quad processor, and everything is running free/open source drivers.

There is only so much code that you can hide inside 58 KB of blobs and that early in the boot sequence, you can't rely on anything else being operational in the device, so you would need to have all the code to initialize and control components on the phone. Think about how much code would be needed to initialize the cellular modem or WiFI and then run a TCP/IP stack to communicate with the outside world. It isn't hard to verify that the blobs that are stored inside the L5's SPI NOR Flash chip are the same ones being distributed by NXP, so then you are left with the theory that NXP or Synopsys are distributing blobs that do something malicious, which would be suicidal for either of those companies if anyone ever discovered it. Supermicro's stock lost 40% of its value after Bloomberg published one story about the Chinese government inserting spy chips in Supermicro motherboards, and nothing in Bloomberg's article was verifiable. Companies like NXP and Synopsys are very unlikely to risk their businesses, even if the NSA asks them, so I find the whole scenario far-fetched.

Re: Intel completely disables AVX-512 on Alder Lake after all

#288
post #254

Earlier quoted context omitted.

I've voiced this before but I think AVX-512 is completely irrelevant in the consumer space (which imho makes titles like "Intel artificially slows 12th gen down" incredulous). Even for commercial applications - sure there are some things that benefit from it. On the other hand, I work in HPC and even in (commercial, not nuke simulations or whatever the nuke powers do with their supercomputers) HPC applications AVX-51…

> HPC applications AVX-512 is rarely used Your applications don't do linear algebra or FFT? AVX-512 is overrated for HPC generally, but it's surely going to be used by an optimized BLAS (unless there's only one FMA unit per core if the implementation is careful). Compute nodes actually should have a big.little structure with a service core for the batch daemon etc. Elsewhere, I see ~130 AVX512 instructions in this De…

Glibc has this strange attraction to having SIMD everything for SIMD friendly functions even though average length to these functions is generally less than 16

Re: Intel completely disables AVX-512 on Alder Lake after all

#289
post #245

Earlier quoted context omitted.

They don't support CPUs that were overclocked, yet advertise that feature, so support isn't the reason.

> yet advertise that feature An therein lies the difference. They advertise the feature, which is something they never did for AVX-512 on Alder Lake desktop. If they advertise a feature, they cannot disable it without getting into legal trouble. If you get K-SKU, you get an unlocked multiplier. That's guaranteed by Intel and that's were their support begins and ends. They cannot do the same for AVX-512, though, becau…

>though, because that's not possible in their heterogeneous architecture

Nobody asks them to enable both AVX-512 and E-cores. People want them to not disable the mode that already works today -- P-cores only mode with AVX-512. There are no OS or driver reasons to not do that.

Re: Intel completely disables AVX-512 on Alder Lake after all

#290
post #256

Earlier quoted context omitted.

> On the other hand, I work in HPC and even in (commercial, not nuke simulations or whatever the nuke powers do with their supercomputers) HPC applications AVX-512 is rarely used. Even in the nuke simulations it is rarely used. More recent cores might be better, but the frequency drop and the associated latency kill performances on the clusters I know. And the new generation ones are AMD anyway.

As ever, it depends, probably on whether your code is dominated by matrix-matrix linear algebra. BLIS DGEMM on my SKX workstation runs at ~88GF, or ~48 if I restrict it to using the haswell configuration (somewhat different on a typical compute node). But yes, I'd rather have twice the cores and memory bandwidth with AVX2. For those that don't know: non-benchmark code usually doesn't get close to peak floating point…

Sure, there are gains if the code is right and the density of AVX instructions is high enough. In most cases that’s not really the case as these dense matrix multiplications are part of a larger algorithm, which includes things like time integration and some calculations that are not straightforward matrix products. The logic is also more complex. The state changes and reduced frequency caused by AVX-512 make the tradeoff difficult.

And as you say, if you have to redesign the code you might as well do it for a GPU that does not have the same issues.

Post reply on HN