Live data from Hacker News

Apple M3 Ultra

apple.com

321–330 of 1001 posts

Re: Apple M3 Ultra

#321

Earlier quoted context omitted.

Transformers are typically memory- bandwidth bound during decoding. This chip is going to have a much worse memory b/w than the nvidia chips. My guess is that these chips could be compute-bound though given how little compute capacity they have.

It's pretty close. A 3090 or 4090 has about 1TB/s of memory bandwidth, while the top Apple chips have a bit over 800GB/s. Where you'll see a big difference is in prompt processing. Without the compute power of a pile of GPUs, chewing through long prompts, code, documents etc is going to be slower.

nobody in industry is using a 4090, they are using H100s which have 3TB/s. Apple also doesn’t have any equivalent to nvlink.

I agree that compute is likely to become the bottleneck for these new Apple chips, given they only have like ~0.1% the number of flops

Re: Apple M3 Ultra

#322

Earlier quoted context omitted.

why would you ever want to do that remains an open question

Probably some kind of local LLM server. 1TB of 1.6 TB/s memory if you link 2 together. $20k total. Half the price of a single Blackwell chip.

with a vanishingly small fraction of flops and a small fraction of memory bandwidth

Re: Apple M3 Ultra

#323

Earlier quoted context omitted.

Framework said that when they built a Strix Halo machine, AMD assigned an engineer to work with them on seeing if there's a way to get CAMM2 memory working with it, and after a bunch of back and forth it was decided that CAMM2 still made the traces too long to maintain proper signal integrity due to the 256 bit interface. These machines have a 512 bit interface, so presumably even worse.

Current (individual, not counting dual socketed) AMD Epyc CPUs have 576 GB/s over a 768 bit bus using socketed DIMMs.

My understanding is that works out due to the lower clock speeds of those RAM modules though right?

It's getting that bandwdith by going very wide on very very very many channels, rather than trying to push a gigantic amount of bandwidth through only a few channels.

Re: Apple M3 Ultra

#324
post #318

Earlier quoted context omitted.

It's the Ultra chip, the same one that goes into the rackmount Mac Pro. I don't think there's much confusion as to who this is for. > there’s no sign that they’re interested in updating that for the AI age. https://security.apple.com/blog/private-cloud-compute/

Outside of extremely niche use cases, who is racking apple products in 2025?

There's MacMiniVault (nee MacMiniColo) https://www.macminivault.com/

Not sure if they count as niche or not.

Re: Apple M3 Ultra

#325
post #222

512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions. I'm curious about how…

They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand. The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new…

Guess what? I'm on a mission to completely max out all 512GB of mem...maybe by running DeepSeek on it. Pure greed!

Re: Apple M3 Ultra

#326

Whoa. M3 instead of M4. I wonder if this was basically binning, but I thought that I had read somewhere that the interposer that enabled this for the M1 chips where not available. That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use.…

If Apple supported Linux (headless) natively, and we could rack m4 pros, I absolutely would use them in our Colo. The CPUs have zero competition in terms of speed, memory bandwidth. Still blown away no other company has been able to produce Arm server chips that can compete.

If I read this right, the r8g.48xlarge at AMZN [1] has 192 cores and 1536GB which exceeds the M3 Ultra in some metrics.

It reminds me of the 1990s when my old school was using Sun machines based on the 68k series and later SPARC and we were blown away with the toaster-sized HP PA RISC machine that was used for student work for all the CS classes.

Then Linux came out and it was clear the 386 trashed them all in terms of value and as we got the 486 and 586 and further generations, the Intel architecture trashed them in every respect.

The story then was that Intel was making more parts than anybody else so nobody else could afford to keep up the investment.

The same is happening with parts for phones and TSMC's manufacturing dominance -- and today with chiplets you can build up things like the M3 Ultra out of smaller parts.

[1] https://aws.amazon.com/ec2/instance-types/r8g/

Re: Apple M3 Ultra

#327
post #61

Let's say you want to have the absolute max memory(512GB) to run AI models and let's say that you are O.K. with plugging a drive to archive your model weights then you can get this for a little bit shy of $10K. What a dream machine. Compared to Nvidia's Project DIGITS which is supposed to cost $3K and be available "soon", you can get a specs matching 128GB & 4TB version of this Mac for about $4700 and the difference…

> I can't wait to see someone testing the full DeepSeek model on this at 819 GB per second bandwidth, the experience would be terrible

Not sure why you are being downvoted, we already know the performance numbers due to memory bandwidth constraints on the M4 Max chips, it would apply here as well.

525GB/s to 1000GB/s will double the TPS at best, which is still quite low for large LLMs.

Re: Apple M3 Ultra

#328
post #139

Earlier quoted context omitted.

I torrent things from two different hosts on my gigabit network. The macos stack literally cannot handle the full bandwidth I have. It fails and the machine needs to be rebooted to fix it. It’s not pretty on the way into this state, either. Other remote connections to the computer are unreliable. On Linux, running the same app in a docker container works perfectly. Transmission is the app.

>Transmission is the app. Former Transmission user here. I realise you didn't ask, but you might find some improvements in qBittorrent.

I went to Transmission years and years ago because it's just simple. It has all the options if you need them, but no HUUUGE interface with RSS feeds, 10001 stats about your download, categories, tags, etc etc etc.

Transmission is just a small, floating window with your downloads. Click for more. It fits in the macOS vibe. But I'm a person that fully adopted the original macOS "way of working" - kicked the full-screen habit I had in windows and never felt better.

Can I ask, why would you go FROM Transmission to qBittorrent?

Re: Apple M3 Ultra

#329
post #22

Previous model of M2 Ultra had max memory of 192GB. Or 128GB for Pro and some other M3 model, which I think is plenty for even 99.9% of professional task. They now bump it to 512GB . Along with insane price tag of $9499 for 512GB Mac Studio. I am pretty sure this is some AI Gold rush.

Maybe .1% of tasks need this RAM, why are they charging so much?

It enables the use of giant AI models on a personal computer. Might not run too fast though. But at least it's possible at all.

Re: Apple M3 Ultra

#330

Earlier quoted context omitted.

Probably never. We don't have official Linux support for the iPhone or iPad, I would't hold out hope for Apple to change their tune.

That makes sense to me though. If you don’t run iOS, you don’t have App Store and that means a loss of revenue.

You lose out on revenue from people who require OS freedom though
Post reply on HN