Live data from Hacker News

AMD's Strix Point: Zen 5 Hits Mobile

chipsandcheese.com

61–70 of 241 posts

Re: AMD's Strix Point: Zen 5 Hits Mobile

#61
post #51
post #48

Earlier quoted context omitted.

Let's do the math on M1 Pro (10-core, N5, 2021) vs HX370 (12-core, N4P, 2024). Firestorm without L3 is 2.281mm2. Icestorm is 0.59mm2. M1 Pro has 8P+2E for a total of 19.428mm2 of cores included. Zen4 without L3 is 3.84mm2. Zen4c reduces that down to 2.48mm2. Zen5 CCD is pretty much the same size as Zen4 (though with 27% more transistors), so core size should be similar. AMD has also stated that Zen5c has a similar sh…

Power efficiency is a curve, and Apple may have its own reason not to make M1 Pro run at 110W as well

I stacked the deck in AMD's favor using a 3-year-old chip on an older node.

Why is AMD using 3.6x more power than M1 to get just 32% higher performance while having 17% more cores? Why are AMD's cores nearly 2x the size despite being on a better node and having 3 more years to work on them?

Why are Apple's scores the same on battery while AMD's scores drop dramatically?

Apple does have a reason not to run at 120w -- it doesn't need to.

Meanwhile, if AMD used the same 33w, nobody would buy their chips because performance would be so incredibly bad.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#62
post #57
post #48

Earlier quoted context omitted.

Let's do the math on M1 Pro (10-core, N5, 2021) vs HX370 (12-core, N4P, 2024). Firestorm without L3 is 2.281mm2. Icestorm is 0.59mm2. M1 Pro has 8P+2E for a total of 19.428mm2 of cores included. Zen4 without L3 is 3.84mm2. Zen4c reduces that down to 2.48mm2. Zen5 CCD is pretty much the same size as Zen4 (though with 27% more transistors), so core size should be similar. AMD has also stated that Zen5c has a similar sh…

119W for hx370 looks extremely sus, seems to me more like the system level power consumption and not CPU-only. According to phoronix [1,2], in their blender CPU test, they measured a peak of 33W. Here max power numbers from some other tests that I know are multi-threaded: -- Linux 6.8 Compilation: 33.13 W LLVM Compilation: 33.25 W -- If I plug in 33W into your equation, that would give us score of HX 370: 104 PPA Thi…

https://www.notebookcheck.net/AMD-Zen-5-Strix-Point-CPU-anal...

They got those kinds of numbers across multiple systems. You can take it up with them I guess.

I didn't even mention one of these systems was peaking at 59w on single-core workloads.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#63
post #58
post #51

Earlier quoted context omitted.

Power efficiency is a curve, and Apple may have its own reason not to make M1 Pro run at 110W as well

I think the OC might have mis-read the power numbers, 110 W is well into desktop CPU power range. Here is a excerpt from Anand Tech: > In our peak power test, the Ryzen AI 9 HX 370 ramped up and peaked at 33 W. https://www.anandtech.com/show/21485/the-amd-ryzen-ai-hx-370...

You can read the notebookcheck review for yourself.

https://www.notebookcheck.net/AMD-Zen-5-Strix-Point-CPU-anal...

Re: AMD's Strix Point: Zen 5 Hits Mobile

#64
post #62
post #57

Earlier quoted context omitted.

119W for hx370 looks extremely sus, seems to me more like the system level power consumption and not CPU-only. According to phoronix [1,2], in their blender CPU test, they measured a peak of 33W. Here max power numbers from some other tests that I know are multi-threaded: -- Linux 6.8 Compilation: 33.13 W LLVM Compilation: 33.25 W -- If I plug in 33W into your equation, that would give us score of HX 370: 104 PPA Thi…

https://www.notebookcheck.net/AMD-Zen-5-Strix-Point-CPU-anal... They got those kinds of numbers across multiple systems. You can take it up with them I guess. I didn't even mention one of these systems was peaking at 59w on single-core workloads.

I see what's going on, they have two HX370 laptops:

  Laptop  MC score  Avg Power
     P16      1213      113 W
     S16       921       29 W
  M3 Pro      1059    (30 W?)
They don't have M3 Pro power numbers, but I assume it is somewhere around 30W, seems like S16 has similar power efficiency as HX 370 at 30 W.

Any more power, and the CPU is much less power efficient, 300% increase in power for 30% increase in performance.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#65
One of these has to be true (or both true):

1. ARM is inherently more efficient than x86 CPUs in most tasks

2. Nuvia and Apple are better CPU designers than AMD and Intel

Here are results from Notebookcheck:

Cinebench R24 ST perf/watt

* M3: 12.7 points/watt

* X Elite: 9.3 points/watt

* AMD HX 370: 3.74 points/watt

* AMD 8845HS: 3.1 points/watt

* Intel 155H: 3.1 points/watt

In ST, Apple is 3.4x more efficient than Zen5. X Elite is 2.4x more efficient than Zen5.

Cinebench R24 MT perf/watt

* M3: 28.3 points/watt

* X Elite: 22.6 points/watt

* AMD HX 370: 19.7 points/watt

* AMD 8845HS: 14.8 points/watt

* Intel 155H: 14.5 points/watt

In MT, Apple is 1.9x more efficient than Zen4 and 1.4x more efficient than HX 370. I expect M3 Pro/Max to increase the gap because generally, more cores means more efficiency for Cinebench MT. X Elite is also more efficient but the gap is closer. However, we should note that in a laptop, ST matters more for efficiency because of the burst behavior of usage. It's easier to gain in MT efficiency as long as you have many cores and run them at lower wattage. In this case, AMD's Zen5 12 core setup and 24 threads works well in Cinebench. Cinebench loves more threads.

One thing that is intriguing is that X Elite does not have little cores which hurts its MT efficiency. It's likely a remnant of Nuvia designing a server CPU, which does not need big.Little but Qualcomm used it in a laptop SoC first.

Sources: https://www.youtube.com/watch?v=ZN2tC8DfJnc

https://www.notebookcheck.net/AMD-Zen-5-Strix-Point-CPU-anal...

Re: AMD's Strix Point: Zen 5 Hits Mobile

#66
>Read bandwidth from a single cluster caps out at just under 62 GB/s. The memory controller has a bit more bandwidth on tap, but you’ll need to load cores from both clusters to get it.

Except for DRR5-7500 it isn't just "a bit more" it is actually double at 120GB/s. This might pose a challenge for LLM inference, which absolutely needs the full 120GB/s.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#67
post #56

I sit firm in my belief that the best thing Microsoft could do for their laptop ecosystem is to add support for a "max fan speed" slider somewhere prominent in the Windows UI. People want the option to make their laptop silent or nearly silent. And when users do need the power, they generally prefer a slightly slower laptop at a reasonable volume rather than the roar of a jet engine. Laptop manufacturers want their d…

I don't know anything about Windows, but at least on Mac, I've been using TGPro for years [0]. I'd assume there is something similar in the Windows world.

In normal conditions my M1 mac can control its fans just fine, but when I travel to hot places like Vietnam... I just keep the fans on more often and my machine doesn't get nearly as hot. I end up having to open it up after a few months and clean out the fans, but that's fine.

[0] https://www.tunabellysoftware.com/tgpro/

Re: AMD's Strix Point: Zen 5 Hits Mobile

#68
post #31

Earlier quoted context omitted.

You missed the part where they mention ARM ends up implementing the same thing to go fast. The point is processors are either slow and efficient, or fast and inefficient. It's just a tradeoff along the curve.

ARM doesn't need the variable-length instruction decoding though, which on x86 essentially means that the decoder has to attempt to decode at every single byte offset for the start of the pipeline, wasting computation. Indeed pretty much any architecture can benefit from some form of op cache, but less of a need for it means its size can be reduced (and savings spent in more useful ways), and you'll still need actual…

x86 processors simply run a instruction length predictor the same way they do it for branch prediction. That turns the problem into something that can be tuned. Instead of having to decode the instruction at every byte offset, you can simply decide to optimize for the 99% case with a slow path for rare combinations.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#69
post #61
post #51

Earlier quoted context omitted.

Power efficiency is a curve, and Apple may have its own reason not to make M1 Pro run at 110W as well

I stacked the deck in AMD's favor using a 3-year-old chip on an older node. Why is AMD using 3.6x more power than M1 to get just 32% higher performance while having 17% more cores? Why are AMD's cores nearly 2x the size despite being on a better node and having 3 more years to work on them? Why are Apple's scores the same on battery while AMD's scores drop dramatically? Apple does have a reason not to run at 120w --…

You should try not to talk so confidently about things you don't know about -- this statement

> if AMD used the same 33w, nobody would buy their chips because performance would be so incredibly bad

Is completely incorrect, as another commenter (and I think the notebookcheck article?) point out -- 30w is about the sweet spot for these processors, and the reason that 110w laptop seems so inefficient is because it's giving the APU 80w of TDP, which is a bit silly since it only performs marginally better than if you gave it e.g. 30 watts. It's not a good idea to take that example as a benchmark for the APU's efficiency, it varies depending on how much TDP you give the processor, and 80w is not a good TDP for these

Re: AMD's Strix Point: Zen 5 Hits Mobile

#70

Earlier quoted context omitted.

[flagged]

What's insane? The claim that recall is spyware isn't that bad of an exaggeration, and the claim that the NPU is hype is a reasonable opinion to have. If that 7% of the die goes unused it's not a huge waste but it's not very good either. If you're not using it enough to affect your battery then you might not get much benefit, because the GPU can do the same tasks about half as fast, it's just less efficient (and it w…

If my speculation about block FP16 is correct, it might be possible to hit 24 to 48 teraflops on the NPU. This means it would be entirely memory bottlenecked even in prompt processing. There would basically be no application where you would run into the NPU being a limitation. What I fear though is that the NPU will be gimped with a single infinity fabric port, which would limit it to a relatively weak 62GB/s out of 120GB/s. For crazy people who want to use 128k token context, that might turn out to make it or break it, in the end.
Post reply on HN