Live data from Hacker News

Apple M1 Ultra

apple.com

451–460 of 855 posts

Re: Apple M1 Ultra

#451
post #215

Earlier quoted context omitted.

My 1st gen 16 core Threadripper is barely faster than an M1 Pro/Max at kernel builds, so a 64 core TR3 should handily double the M1 Ultra performance. But you know, I'm still happy to double my current build perf in a small box I can stick in my closet. Ordered one :-)

How many threads are actually getting utilized in those kernel builds? I don't work on the kernel enough to have intuition in mind but people make wildly optimistic assumptions about how compilation stresses processors. Also 1st gen threadrippers are getting on a bit now, surely. It's a ~6 year old microarchitecture.

Kernel compilation can be heavily parallelized.

Re: Apple M1 Ultra

#452
post #368

Earlier quoted context omitted.

Next-gen Threadripper Pro was also announced today: https://www.tomshardware.com/news/amd-details-ryzen-threadri...

Bit of a wet fart though, even Charlie D thinks it's too little too late. OEM-only (and only on WRX80 socket), no V-cache, worse product support. https://semiaccurate.com/2022/03/08/amd-finally-launches-thr... The niche for high clocks was arguable with the 2nd-gen products but now you are foregoing v-cache which also improves per-thread performance, so Epyc is relatively speaking even more attractive. And if you tak…

The PRO Threadrippers are not cut down, they have 8 memory channels and 128 PCIE 4.0 lanes. I think the only limitation compared to Epyc is that you can have 1 socket only.

Re: Apple M1 Ultra

#453

I think the GPU claims are interesting. According to the graph's footer, the M1 Ultra was compared to an RTX 3090. If the performance/wattage claims are correct, I'm wondering if the Mac Studio could become an "affordable" personal machine learning workstation (which also won't make the electricity bill skyrocket). If Pytorch becomes stable and easy to use on Apple Silicon [0][1], it could be an appealing choice. [0]…

I know Nvidia has never really care much about TDP, but this still seems unbelievable to me. How could a relatively new design beat a 3090 with 200W less power, while having to share a die with a CPU? It just doesn't seem possible.

Re: Apple M1 Ultra

#454

Earlier quoted context omitted.

Could they do a multi-socket board for the Mac Pro?

They would never do that

The Mac Pro (and the high-end Powermacs that preceded it) were always available in a dual socket incarnation, right up to the trashcan Mac.

Re: Apple M1 Ultra

#455
post #230

Is it actually more powerful than a top of the line threadripper[0] or is that not a "personal computer" CPU by this definition? I feel like 64 cores would beat 20 on some workloads even if the 20 were way faster in single core performance. [0] https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-3...

My workstation has a 3990x. Our "world" build is slightly faster on my M1 Max. https://twitter.com/kiratpandya/status/1457438725680480257 The 3990x runs a bit faster on the initial compile stage but the linking is single threaded and the M1 Max catches up at that point. I expect the M1 Ultra to crush the 3990x on compile time.

From the tweet reply you used sccache with hot cache, which would probably be mostly single-threaded since it's just fetching and copying things from cache.

Re: Apple M1 Ultra

#456

Curious about the naming here, terms like "pro" "pro max" "max", "ultra" (hopefully there's no "pro ultra" and "ultra max" in the future) is very confusing and hard to know which one is more powerful than which, or if it's a power-level relationship. Is this on purpose or it's just bad naming? Is there example of good naming for this kind of situation?

I think the naming is based entirely around how the announcement sentence lands in the keynote. So they've optimized for "We're adding one last chip to the M1 family, and it's gonna blow your mind... (M1 Ultra appears on screen)".

Re: Apple M1 Ultra

#457
post #35

Earlier quoted context omitted.

>I wonder if they'll just do 4x M1 Max for that. They'll be running out of names for that thing. M1 Ultra II would be lame, so M1 Extreme? M1 Steve?

M1 Plaid

lol, good one! I hope marketing sees this.

Re: Apple M1 Ultra

#459
post #432
post #353

Earlier quoted context omitted.

AMD will be on TSMC N5P next year, which will give them node parity with Apple (who will be releasing A15 on N5P this year), and actually a small node lead over the current N5-based A14 products. So we will get to test the "it's all just node lead guys, nothing wrong with x86!!!" theory. Don't worry though there will still be room to move the goalposts with "uhhh, but, Apple is designing for high IPC and low clocks,…

We've already seen x86 draw even with Intel 12th gen: https://www.youtube.com/watch?v=X0bsjUMz3EM

So in that one, you've got a 20-thread Intel part (6+8C/20T) at probably 3.5 GHz going against a 10-thread Apple part (8+2C/10T), at probably 3 GHz, and the Apple part still comes out on top by ~5% in Cinebench R23 MT. And that's with Intel having 50% more high-performance threads available.

Work out the IPC there - the Intel has a 2x thread count advantage, a 17% clock advantage, and Apple comes out 5% ahead. So the IPC gap there is about 2.46x.

It's not a perfect comparison of course, since we're mixing SMT and big/little cores, but in basically every area Intel should (on paper) have more resources available and Apple is coming out on top anyway by sheer IPC.

That's what I'm saying - you can't really do that approach with x86. It's not power-advantageous or transistor-advantageous to go super wide on the decode or reorder buffer like that on x86. And regardless of the tricks x86 uses to mitigate it, you've still got a 2.5x IPC gap at the end of the day. A 2.5x IPC gap will not be closed up by just a single node shrink.

And that's looking at MT, where your task scales perfectly. See where I'm going with this? Intel is using 2x the number of threads, and 3x the number of efficiency cores to get there. Apple can deliver that punch across a much lower number of threads - meaning ST-bottlenecked tasks will scale much much better on Apple.

With a single-threaded test, the M1 is pulling 7W vs 33W for the Alder Lake intel. Obviously that tells us nothing about efficiency, since we'd need to know the scores, but that's the downside, is for normal, poorly-threaded tasks, like surfing the web or editing code, the 12900HK is going to be boosting high to reach the same performance levels the M1 does at 3 GHz. And that's exactly what you see in the power figures there.

In short: you will likely see x86 able to keep up in one metric or another. You can win on performance if you just go nuclear on power. You can match on power on perfectly-threadable tasks that allow the x86 to deploy twice the threads (sharing instruction cache/etc). You can match on single-threaded battery life if you accept lesser performance. But the overall performance of the M1 derives from the massive IPC it generates, and that's something that x86 can't match nearly as easily.

Going ham on a single metric just to claim victory isn't nearly the same thing as the level of all-round performance and efficiency that Apple has achieved there.

(see also, putting a 128-thread Threadripper 3990WX workstation up against a 10-thread M1 Max laptop just to win at rendering... and people here thought that disproved that Apple was great hardware lol)

Re: Apple M1 Ultra

#460

Earlier quoted context omitted.

I've been saying 4x M1 Max is not a thing and never will be a thing ever since the week I got my M1 Max and saw that the IRQ controller was only instantiated to support 2 dies, but everyone kept parroting that nonsense the Bloomberg reporter said about a 4-die version regardless... Turns out I was right. The Mac Pro chip will be a different thing/die.

Plus they are running out of M1 superlatives. They’ll have to go to M2 to avoid launching M1 Plaid.

m1 hyper turbo deluxe

or they could take a page out of microsofts book and just call the next one "m one"

Post reply on HN