Live data from Hacker News

AMD's Strix Point: Zen 5 Hits Mobile

chipsandcheese.com

201–210 of 241 posts

Re: AMD's Strix Point: Zen 5 Hits Mobile

#201
post #77
post #48

Earlier quoted context omitted.

Let's do the math on M1 Pro (10-core, N5, 2021) vs HX370 (12-core, N4P, 2024). Firestorm without L3 is 2.281mm2. Icestorm is 0.59mm2. M1 Pro has 8P+2E for a total of 19.428mm2 of cores included. Zen4 without L3 is 3.84mm2. Zen4c reduces that down to 2.48mm2. Zen5 CCD is pretty much the same size as Zen4 (though with 27% more transistors), so core size should be similar. AMD has also stated that Zen5c has a similar sh…

Even with the M3 the difference is marginal in multi-threaded benchmarks, from the Cinebench link [1] someone posted earlier on the thread. Apple M3 Pro 11-Core - 394 Points per Watt AMD Ryzen AI 9 HX 370 - 354 Points per Watt Apple M3 Max 16-Core - 306 Points per Watt And the Ryzen in on TSMC 4nm while the M3 is on 3nm. As parent is saying, a lot of the Apple Silicon hype was due to the massive upgrade it was over t…

Their efficiency tests use Cinebench R23 (as called out explicitly).

R23 is not optimized for Apple silicon but is for x86. The R24 numbers are actually what you need for a fair comparison, otherwise you put the Arm numbers at a significant handicap.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#202

Earlier quoted context omitted.

>The third possibility is that they just pick a different point on the efficiency curve. You can double power consumption in exchange for a few percent higher performance, double it again for an even smaller increase. This only makes sense if the Zen5 is actually faster in ST than the M3. In this case, the M3 is 1.24x faster and 3.4x more efficient in ST than Zen5. AMD's Zen5 chip is just straight up slower in any cu…

> This only makes sense if the Zen5 is actually faster in ST than the M3. In this case, the M3 is 1.24x faster and 3.4x more efficient in ST than Zen5. It makes sense if Zen5 is faster in MT, since that's when the CPUs will be power limited, and it is. For ST the performance generally isn't power-limited for either of them and then the M3 is on a newer process node. It also depends on the benchmark. For example, Zen5…

R23 doesn’t have the same SIMD optimizations available for ARM as it does for x86.

Anyone using R23 instead of R24, is putting arm at a disadvantage. Notebookcheck is often called out for this and haven’t really addressed why they stick with R23 beyond not wanting to redo tests for older hardware. They are by far the outlier for performance numbers and why the discussion around performance gets muddied.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#203

Earlier quoted context omitted.

>I sit firm in my belief that the best thing Microsoft could do for their laptop ecosystem is to add support for a "max fan speed" slider somewhere prominent in the Windows UI Until then, there's https://github.com/Rem0o/FanControl.Releases

A closed source application for controlling one's fan...umm no thank you. I never will understand the reasoning behind why people are so afraid of releasing their source code. Looks like a weekend project; does he expect to make a living out of a weekend project?

In this case, every OEM will just copy it and slap their name on it. He released it as freeware and he has every right to do so. You have no right to others work.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#204
post #202

Earlier quoted context omitted.

> This only makes sense if the Zen5 is actually faster in ST than the M3. In this case, the M3 is 1.24x faster and 3.4x more efficient in ST than Zen5. It makes sense if Zen5 is faster in MT, since that's when the CPUs will be power limited, and it is. For ST the performance generally isn't power-limited for either of them and then the M3 is on a newer process node. It also depends on the benchmark. For example, Zen5…

R23 doesn’t have the same SIMD optimizations available for ARM as it does for x86. Anyone using R23 instead of R24, is putting arm at a disadvantage. Notebookcheck is often called out for this and haven’t really addressed why they stick with R23 beyond not wanting to redo tests for older hardware. They are by far the outlier for performance numbers and why the discussion around performance gets muddied.

> R23 doesn’t have the same SIMD optimizations available for ARM as it does for x86.

The single-thread benchmark is SIMD-heavy?

Now it just sounds like Cinebench ST is a useless benchmark because it's putting a parallelizable SIMD workload on a single core. In real life you'd always be running those multi-threaded, whereas the reason people care about ST performance is for the serialized branch-heavy spaghetti code that inherently only runs on one core. "Run the SIMD code, but clamp it to a single thread" is a garbage proxy for that.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#205
post #140

Earlier quoted context omitted.

> I stacked the deck in AMD's favor using a 3-year-old chip on an older node. You could just compare the ones that are actually on the same process node: https://www.notebookcheck.net/R9-7945HX3D-vs-M2-Max_15073_14... But then you would see an AMD CPU with a lower TDP getting higher benchmark results. > Why is AMD using 3.6x more power than M1 to get just 32% higher performance while having 17% more cores? Getting 32…

> Getting 32% higher performance from 17% more cores implies higher performance per core. I don't disagree that it is higher perf/core. It is simply MUCH worse perf/watt because they are forced to clock so high to achieve those results. > The power measurements that site uses are from the plug, which is highly variable to the point of uselessness They measure the HX370 using 119w with the screen off (using an externa…

> I don't disagree that it is higher perf/core. It is simply MUCH worse perf/watt because they are forced to clock so high to achieve those results.

The base clock for that CPU is nominally 2 GHz.

> They measure the HX370 using 119w with the screen off (using an external monitor). What on that motherboard would be using the remaining 85+W of power?

For the Asus ProArt P16 H7606WI? Probably the 115W RTX 4070.

> TDP is a suggestion, not a hard limit. Before thermal throttling, they will often exceed the TDP by a factor of 2x or more.

TDP is not really a suggestion. There are systems that can't dissipate more than a specific amount of heat and producing more than that could fry other components in the system even if the CPU itself isn't over-temperature yet, e.g. because the other components have a lower heat tolerance. There are also systems that can't supply more than a specific amount of power and if the CPU tried to non-trivially exceed that limit the system would crash.

The TDP is, however, configurable, including different values for boost. So if the OEM sets the value to the higher end of the range even though their cooling solution can't handle it, the CPU will start out there and gradually lower its power use as it becomes thermally limited. This is not the same as "TDP is a suggestion", it's just not quite as simple as a single number.

> As to these specific benchmarks, the R9 7945HX3D you linked to used 187w while the M2 Max used 78w for CB R15.

Which is the same site measuring power consumption at the plug on an arbitrary system with arbitrary other components drawing power. Are they even measuring it though the power brick and adding its conversion losses?

These CPUs have internal power meters. Doing it the way they're doing it is meaningless and unnecessary.

> You should be looking at benchmarks without such a massive bias.

Do you have one that compares the same CPUs on some representative set of tests and actually measures the power consumption of the CPU itself? Diligently-conducted benchmarks are unfortunately rare.

Note however that the same link shows the 7945HX3D also ahead in Blender, Geekbench ST and MT, Kraken, Octane, etc. It's consistently faster on the same process, and has a lower TDP.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#206
post #202

Earlier quoted context omitted.

R23 doesn’t have the same SIMD optimizations available for ARM as it does for x86. Anyone using R23 instead of R24, is putting arm at a disadvantage. Notebookcheck is often called out for this and haven’t really addressed why they stick with R23 beyond not wanting to redo tests for older hardware. They are by far the outlier for performance numbers and why the discussion around performance gets muddied.

> R23 doesn’t have the same SIMD optimizations available for ARM as it does for x86. The single-thread benchmark is SIMD-heavy? Now it just sounds like Cinebench ST is a useless benchmark because it's putting a parallelizable SIMD workload on a single core. In real life you'd always be running those multi-threaded, whereas the reason people care about ST performance is for the serialized branch-heavy spaghetti code t…

Yes, there’s no difference between the single and multi core benchmark other than how many threads get spun up.

I’m not sure why you’re trying to equate simd with parallelization. Tbh, a lot of your response seems odd to me because it’s making several incorrect assumptions.

You can’t really escape parallelization with how any modern core works, even on a single core. You may have certain operations process concurrently depending on how the cores resources are available at a given time and what is needed.

Regardless, SIMD isn’t concurrency. It’s batching.

There’s still significant benefit to having SIMD on a single threaded task. There’s a lot of thread overhead to using multiple cores to do something, whereas SIMD lets you effectively batch things on a single core.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#207

Earlier quoted context omitted.

"Thin and fanless" aren't that hard, just use any low power CPU. But then people also want fast. Apple does this by buying out TSMC's capacity for the latest process nodes and then taking the performance/efficiency trade off in favor of efficiency, so they get something with similar performance and lower power consumption. But then they charge you $400 for $50 worth of RAM and solder it so you can't upgrade it yourse…

I know everyone on this site loves to hate on soldered ram, but my impression is most people don’t understand that soldered ram is not the same thing as regular ram modules. They are literally different memory chips (LPDDR vs DDR) . When built to a specific chip my understanding is you can design for tighter timings and higher bandwidth which is important for the gpu. The M1 shipped with very fast LPDDR4X running at…

The objection to soldered RAM isn’t the RAM.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#208

Earlier quoted context omitted.

Fast is relative. The Ryzen HX 370 has a TDP configurable down to 15W and at that power level it could be run fanless and would be faster than the M1, but it's still faster yet if you give it 54W and raise the clock speed.

I'm going to need source on that. What does HX 370 score at 10w?

You're asking for a benchmark result for a CPU which just came out and has a configurable TDP that hardly anybody is going to have set to its lowest value, if they even disclose it, much less have done so in a test against the original M1. If you think a source for that even exists you can provide a link.

But the result seems pretty obvious. Even the 7nm Ryzen U-series at 15W (e.g. 7730U) was beating the 5nm M1 on multi-threaded workloads and the HX 370 is well ahead of both on single-thread performance. Single-thread workloads aren't significantly power limited, so to not be the case the Zen5 HX 370 would have to be slower than the Zen3 7730U on threaded workloads at the same TDP, which seems unlikely.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#209

Earlier quoted context omitted.

Fast is relative. The Ryzen HX 370 has a TDP configurable down to 15W and at that power level it could be run fanless and would be faster than the M1, but it's still faster yet if you give it 54W and raise the clock speed.

Is that the chip AMD just released? Isn't the M1 about 4 years old?

The premise is that others can now use the same process as the M1 did to make fanless CPUs. Which they can, but they could always make fanless CPUs. The issue is that people also want them to be fast, which is not an absolute measurement fixed for all time, it's relative to competing contemporary systems with more cooling, which will always be faster.

Re: AMD's Strix Point: Zen 5 Hits Mobile

#210

Earlier quoted context omitted.

"Thin and fanless" aren't that hard, just use any low power CPU. But then people also want fast. Apple does this by buying out TSMC's capacity for the latest process nodes and then taking the performance/efficiency trade off in favor of efficiency, so they get something with similar performance and lower power consumption. But then they charge you $400 for $50 worth of RAM and solder it so you can't upgrade it yourse…

I know everyone on this site loves to hate on soldered ram, but my impression is most people don’t understand that soldered ram is not the same thing as regular ram modules. They are literally different memory chips (LPDDR vs DDR) . When built to a specific chip my understanding is you can design for tighter timings and higher bandwidth which is important for the gpu. The M1 shipped with very fast LPDDR4X running at…

You're making two separate points here.

The first is the timings, which is nominally real but it was never a huge difference. Moreover, the new CAMM standard aims to address this and basically does. The legacy SODIMM standard wasn't great and is essentially what caused this.

The second is the bus width. If you use slotted memory and want a wide bus then you need at least one slot per channel and then you could end up needing a lot of slots. This isn't impossible -- servers do it -- but there is a cost attached to it.

But this doesn't apply to systems that aren't using a wide bus. The base M3 has the same bus width as ordinary dual-channel PCs. The Pro has the equivalent of four channels or, for the newer generation, three. That's still not a crazy number in a high end system. With CAMM it would only be two modules, since the modules are each 128-bit. By way of comparison, Threadripper has four or eight channels and modern servers have dozens.

Post reply on HN