Live data from Hacker News

Apple M3 Ultra

apple.com

481–490 of 1001 posts

Re: Apple M3 Ultra

#481

Earlier quoted context omitted.

There isn't anything particularly high-bandwidth about Apple's DDR5 implementation, either. They just have a lot of channels, which is why I compared it to a 24-channel EPYC system. I agree that their integrated GPU architecture hits a unique design point that you don't get from nvidia, who prefer to ship smaller amounts of very different kinds of memory. Apple's architecture may be more suited to some workloads but…

M3 Ultra has 819GB/s, and a single epyc cpu with 12 channels has 460GB/s. As far as I know, llama.cpp and friends don’t scale across multiple sockets so you can’t use a dual socket Turin system to match the M3 Ultra. Also, 32GB DDR5 RDIMMS are ~200, so that’s 5K for 24 right there. Then you need 2x CPUs, at ~1K for the cheapest, and you need 2, and then a motherboard that’s another 1K. So for 8K (more, given you need…

Partial correction, an Epyc CPU with 12 channels has 576 GB/s, i.e. DDR5-6000 x 768 bits. That is 70% of the Apple memory bandwidth, but with possibly much more memory (768 GB in your example).

You do not need 2 CPUs. If however you use 2 CPUs, then the memory bandwidth doubles, to 1152 GB/s, exceeding Apple by 40% in memory bandwidth. The cost of the memory would be about the same, by using 16 GB modules, but the MB would be more expensive and the second CPU would add to the price.

Re: Apple M3 Ultra

#482
post #46

They update the Studio to M3 Ultra now, so M4 Ultra can presumably go directly into the Mac Pro at WWDC? Interesting timing. Maybe they'll change the form factor of the Mac Pro, too? Additionally, I would assume this is a very low-volume product, so it being on N3B isn't a dealbreaker. At the same time, these chips must be very expensive to make, so tying them with luxury-priced RAM makes some kind of sense.

Honestly I don't think we'll see the M4 Ultra at all this year. That they introduced the Studio with an M3 Ultra tells me M4 Ultras are too costly or they don't have capacity to build them.

And anyway, I think the M2 Mac Pro was Apple asking customers "hey, can you do anything interesting with these PCIe slots? because we can't think of anything outside of connectivity expansion really"

RIP Mac Pro unless they redesign Apple Silicon to allow for upgradeable GPUs.

Re: Apple M3 Ultra

#483
post #218

Earlier quoted context omitted.

This is more about "average" end user software, not the type of software that would be running on a machine like this. Yes their applications fell off, but if you're paying for 512gb of RAM apple notes being slow isn't the bottleneck

Lack of focus on quality of software affects all types of workloads, not just consumer-oriented or professional-oriented in isolation.

> Lack of focus on quality of software affects all types of workloads, not just consumer-oriented or professional-oriented in isolation.

The apps are developed by different teams. MacOS apps are containerized. Saying macOS's performance is hindered by Notes.app is like saying that Windows is hindered by Paint.exe. Notes.app is just a default[0]

[0]: though, I dislike saying this because I always feel like I need to mention that even Notes links against a hilarious amount of private APIs that could easily be exposed to other developers but... aren't.

Re: Apple M3 Ultra

#484
post #198

Earlier quoted context omitted.

Face ID, taking pictures, Siri, ARKit, voice-to-text transcription, face recognition and OCR in photos, noise filtering, ...

These have been possible in much smaller smartphone chips for years.

Possible != energy efficient, which is important for mobile devices.

Re: Apple M3 Ultra

#485
post #349

Earlier quoted context omitted.

The M4 Pro is 56% faster in ST performance against AMD’s new Strix Halo while being 3.6x more efficient. Source: https://www.notebookcheck.net/AMD-Ryzen-AI-Max-395-Analysis-... Cinebench 2024 results.

That’s a laptop part, so it makes different tradeoffs. Somewhere on the internet there is a tdp wattage vs performance x-y plot. There’s a pareto optimal region where all the apple and amd parts live. Apple owns low tdp, AMD owns high tdp. They duke it out in the middle. Intel is nowhere close to the line. I’d guess someone has made one that includes datacenter ARM, but I’ve never seen it.

> tdp wattage vs performance x-y plot

This?

https://www.videocardbenchmark.net/power_performance.html#sc...

Re: Apple M3 Ultra

#486
That's all nice, but if they are to be considered a serious AI hardware player, they will need to invest in better support of their hardware in deep learning frameworks such as PyTorch and Jax. Currently the support is rather poor, and is not suitable for any serious work.

Re: Apple M3 Ultra

#487
post #349

Earlier quoted context omitted.

That’s a laptop part, so it makes different tradeoffs. Somewhere on the internet there is a tdp wattage vs performance x-y plot. There’s a pareto optimal region where all the apple and amd parts live. Apple owns low tdp, AMD owns high tdp. They duke it out in the middle. Intel is nowhere close to the line. I’d guess someone has made one that includes datacenter ARM, but I’ve never seen it.

High TDP? You mean server-grade CPUs? Apple doesn't make those.

> You mean server-grade CPUs? Apple doesn't make those.

Right.

It is coming up because we're in a thread about using them as server CPUs. (c.f. "colo", "2U" in OP and OP's child), and the person you're replying to is making the same point you are

For years now, people will comment "these are the best chips, I'd replace all chips with them."

Then someone points out perf/watt is not perf.

Then someone else points out some M-series is much faster than a random CPU.

And someone else points out that the random CPU is not a top performing CPU.

And someone else points out M-series are optimized for perf/watt and it'd suck if it wasn't.

I love my MacBook, the M-series has no competitors in the case it's designed for.

I'd just prefer, at this point, that we can skip long threads rehashing it.

It's a great chip. It's not the fastest, and it's better for that. We want perf/watt in our mobile devices. There's fundamental, well-understood, engineering tradeoffs that imply being great at that necessitates the existence of faster processors.

Re: Apple M3 Ultra

#488
post #476
post #459

I know it's basically nitpicking competing luxury sports cars at this point, but I am very bothered that existing benchmarks for the M3 show single core perf that is approximately 70% of M4 single core perf. I feel like I should be able to spend all my money to both get the fastest single core performance AND all the cores and available memory, but Apple has decided that we need to downgrade to "go wide". Annoying.

> both get the fastest single core performance AND all the cores I'm a major Apple skeptic myself, but hasn't there always been a tradeoff between "fastest single core" vs "lots of cores" (and thus best multicore)? For instance, I remember when you could buy an iMac with an i9 or whatever, with a higher clock speed and faster single core, or you could buy an iMac Pro with a Xeon with more cores, but the iMac (non-Pro…

> hasn't there always been a tradeoff between "fastest single core" vs "lots of cores" (and thus best multicore)?

Not in the Apple Silicon line. The M2 Ultra has the same single core performance as the M2 Max and Pro. No benchmarks for the M3 Ultra yet but I'm guessing the same vs M3 Max and Pro.

Re: Apple M3 Ultra

#489
post #43

Too bad it lacks even the streaming mode SVE2 found in M4 cores. If only Apple would provide a full SVE2 implementation to put pressure on ARM to make it non-optional so AArch64 isn't effectively restricted to NEON for SIMD.

Hell I’m just sitting here hoping the future M5 adopts SVE. Not even SVE2.

Re: Apple M3 Ultra

#490
post #422

Earlier quoted context omitted.

It is much slower than nVidia, but for a lot of personal-use LLM scenarios, it's very workable. And it doesn't need to be anywhere near as fast considering it's really the only viable (affordable) option for private, local inference, besides building a server like this, which is no faster: https://news.ycombinator.com/item?id=42897205

It's fast enough for me to cancel monthly AI services on a mac mini m4 max.

Could you maybe share a lightweight benchmark where you share the exact model (+ quantization if you're using that) + runtime + used settings and how much tokens/second you're getting? Or just like a log of the entire run with the stats, if you're using something like llama.cpp, LMDesktop or ollama?

Also, would be neat if you could say what AI services you were subscribed to, there is a huge difference between paid Claude subscription and the OpenAI Pro subscription for example, both in terms of cost and the quality of responses.

Post reply on HN