Live data from Hacker News

Apple M3 Ultra

apple.com

411–420 of 1001 posts

Re: Apple M3 Ultra

#411
post #394

Earlier quoted context omitted.

High TDP? You mean server-grade CPUs? Apple doesn't make those.

True, but these "Ultra" chips do target the same niche as (some) high-TDP chips. Workstations (like the Mac Studio) have traditionally been a space where "enthusiast"-grade consumer parts (think Threadripper) and actual server parts competed. The owner of a workstation didn't usually care about their machine's TDP; they just cared that it could chew through their workloads as quickly as possible. But, unlike an actua…

Oh you. mean Threadripper. I thought you were talking about Epyc.

Anyway, I don't think it's comparable really. This thing comes with a fat GPU, NPU, and unified memory. Threadripper is just a CPU.

Re: Apple M3 Ultra

#412

Earlier quoted context omitted.

That makes sense to me though. If you don’t run iOS, you don’t have App Store and that means a loss of revenue.

You lose out on revenue from people who require OS freedom though

All seven of them. I kid, I have a lot of sympathy for that position, but as a practical matter running Linux VMs on an M4 works great, you even get GPU acceleration.

Re: Apple M3 Ultra

#413

Earlier quoted context omitted.

VRAM is what takes a model from "can not run at all" to "can run" (even if slowly), hence the emphasis.

No, with limited VRAM you could offload the model partially or split across CPU and GPU. And since CPU has swap, you could run the absolute largest model. It’s just really really slow.

The difference between Deepseek-r1:70b (edit: actually 32b) running on an M4 Pro (48 GB unified RAM, 14 CPU cores, 20 GPU cores) and on an AMD box (64 GB DDR4, 16 core 5950X, RTX 3080 with 10 GB of RAM) is more than a factor of 2.

The M4 pro was able to answer the test prompt twice--once on battery and once on mains power--before the AMD box was able to finish processing.

The M4's prompt parsing took significantly longer, but token generation was significantly faster.

Having the memory to the cores that matter makes a big difference.

Re: Apple M3 Ultra

#414
post #402

Earlier quoted context omitted.

High TDP? You mean server-grade CPUs? Apple doesn't make those.

Isn't the rack-mounted Mac Pro supposedly "server-grade" ( https://www.apple.com/shop/buy-mac/mac-pro/rack )? At least judging by the mounts, they want them to be used that way, even though the CPU might not fit with the de facto industry label for "server-grade".

Server grade CPUs. I thought he was referring to Epyc CPUs.

Re: Apple M3 Ultra

#415
post #307

Earlier quoted context omitted.

No, I'm not. I'm comparing the TOPS of the M3 Ultra and the tensor cores of the RTX 5090. If not, what is the TOPS of the GPU, and why isn't apple talking about it if there is more performance hidden somewhere? Apple states 18 TOPS for the M3 Max. And why do you think Apple added the neural engine, if not to accelerate compute? The power draw is quite a bit higher, but it's still much more efficient as the performanc…

The ANE and tensor cores are not comparable though. One is literally meant for low cost inference while the others are meant for acceleration of training. If you squint, yeah they look the same, but so does the microcontroller on the GPU and a full blown CPU. They’re fundamentally different purposes, architectures and scale of use. The ANE can’t even really be used directly. Apple heavily restricts the use via CoreML…

So now the TOPS are not comparable because M3 is much slower than an Nvidia GPU? That's not how comparisons work.

My numbers are correct, the M3 Ultra has around 1 % of the TOPS performance of a RTX 5090.

Comparing against the GPU would look even worse for apple. Do you think Apple added the neural engine just for fun? This is exactly what the neural engine is there for.

Re: Apple M3 Ultra

#416
post #349

Earlier quoted context omitted.

That’s a laptop part, so it makes different tradeoffs. Somewhere on the internet there is a tdp wattage vs performance x-y plot. There’s a pareto optimal region where all the apple and amd parts live. Apple owns low tdp, AMD owns high tdp. They duke it out in the middle. Intel is nowhere close to the line. I’d guess someone has made one that includes datacenter ARM, but I’ve never seen it.

High TDP? You mean server-grade CPUs? Apple doesn't make those.

Indeed. The M3 Ultra is in the midrange where they duke it out. Similarly, for its niche, the iPhone CPU is was better than AMD’s low end processors.

Anyway the Apple config in the article costs about 5x more than a comparable low end AMD server with 512GB of ram, but adds an NPU. AMD has NPUs in lower end stuff; not sure about this TDP range.

Re: Apple M3 Ultra

#417
post #222

512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions. I'm curious about how…

Why does it matter if you can run the LLM locally, if you're still running it on someone else's locked down computing platform?

Re: Apple M3 Ultra

#419
post #285

Earlier quoted context omitted.

For enterprise markets, this is table stakes. A lot of datacenter customers will probably ignore this release altogether since there isn't a high-bandwidth option for systems interconnect.

Thunderbolt 5 can do bi-directional 80 Gbps....and Mac Studio Ultra has 6 ports...

That's still not even competitive with 100G Ethernet on a per-port basis. An overall bandwidth of 480 Gbps pales in comparison with, for example, the 3200 Gbps you get with a P5 instance on EC2.

Re: Apple M3 Ultra

#420
post #382
post #375

Earlier quoted context omitted.

It will cost 4X what it costs to get 512GB on an x86 server motherboard.

What would it cost to get 512GB of VRAM on an Nvidia card? That’s the real comparison.

Since the GH200 has over a terabyte of VRAM at $343,000 and the H100 has 80GB that makes that $195,993 with a bit over 512GB of VRAM . You could beat the price of the Apple M3 Ultra with an AMD EPYC build.
Post reply on HN