They update the Studio to M3 Ultra now, so M4 Ultra can presumably go directly into the Mac Pro at WWDC? Interesting timing. Maybe they'll change the form factor of the Mac Pro, too? Additionally, I would assume this is a very low-volume product, so it being on N3B isn't a dealbreaker. At the same time, these chips must be very expensive to make, so tying them with luxury-priced RAM makes some kind of sense.
Interestingly, Apple apparently confirmed to a French website that M4 lacks the interconnect required to make an "Ultra" [0][1], so contrary to what I originally thought, they maybe won't make this after all? I'll take this report with a grain of salt, but apparently it's coming directly from Apple. Makes it even more puzzling what they are doing with the M2 Mac Pro. [0] https://www.numerama.com/tech/1919213-m4-max-e…
Apple M3 Ultra
801–810 of 1001 posts
Re: Apple M3 Ultra
#802Earlier quoted context omitted.
3.5 - 4.5 tokens/s on the $2,000 AMD Epyc setup. Deepseek 671b q4. The AMD Epyc build is severely bandwidth and compute constrained. ~40 tokens/s on M3 Ultra 512GB by my calculation.
Thanks! If the M3 can run 24/7 without overheating it's a great deal to run agents. Especially considering that it should run only using 350W... so roughly $50/mo in electricity costs.
I'd assume this thing peaks at 350W (or whatever) but idles at around 40w tops?
Re: Apple M3 Ultra
#803Earlier quoted context omitted.
Possible != energy efficient, which is important for mobile devices.
If the energy efficiency of things like Face ID was indeed so far so bad that you need a more efficient M3 Ultra, how come Face ID was integrated into smartphones years ago, apparently without significant negative impact on battery life?
Image recognition, OCR, AR and more are applications of the NPU that didn’t exist at all on older iPhones because they have would be too intensive for the chips and batteries.
Re: Apple M3 Ultra
#804How do people feel about the value of the M3 Ultra vs. the M4 Max for general computing, assuming that you max out the RAM on the M4 version of the Studio?
The kinds of workloads that could truly leverage the M2 Ultra over the M2 Max were vanishingly small. When comparing the M3 Ultra to the M4 Max, that number gets even smaller, because the M4 Max will have ~15% higher single core perf. The insane memory available on M3 Ultra is its only interesting capability, but its still not big enough to run the series of largest open source LLMs. Hot take: You can tie yourself in…
Re: Apple M3 Ultra
#805Earlier quoted context omitted.
M1 came out before the LLM rush, though
The M1 is in a product segment where discrete GPUs have been gone for decades , in favor of integrated graphics that shares one pool of RAM with the CPU. The better question to ask is why Apple kept using that unified memory design even when moving up to larger chips like the M1 Max and M1 Ultra.
So if you wanted to give it a second ram pool you would have to add an entire second memory interface just for the on-die GPU.
Now all you’ve done is make it more complicated, slower because now you have to move things between the two pools, and gained what exactly?
I think it was a very clear and obvious decision to make. It’s an outgrowth out of how the base chips were designed, and it turned out to be extremely handy for some things. Plus since all their modern devices now work this way that probably simplify the software.
I’m not saying it’s genius foresight, but it certainly worked out rather well. There’s nothing stopping them from supporting discreet GPUs too if they wanted to. They just clearly don’t.
Re: Apple M3 Ultra
#806512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions. I'm curious about how…
Is this on chip memory? From the 800GB/s I would guess more likely a 512bit bus (8 channel) to DDR5 modules. Doing it on a quad channel would just about be possible, but really be pushing the envelope. Still a nice thing. As for practicality, which mainstream applications would benefit from this much memory paired with a nice but relative mid compute? At this price-point (14K for a full specced system), would you pre…
Re: Apple M3 Ultra
#807Earlier quoted context omitted.
I think the other big thing is that the base model finally starts at a normal amount of memory for a production machine. You can't get less than 96GB. Although an extra $4000 for the 512GB model seems Tim Apple levels of ridiculous. There is absolutely no way that the different costs anywhere near that much at the fab. And the storage solution still makes no sense of course, a machine like this should start at 4TB fo…
> There is absolutely no way that the different costs anywhere near that much at the fab. price premium probably, but chip lithography errors (thus, yields) at the huge memory density might be partially driving up the cost for huge memory.
Apple absolutely loves to gouge for upgrades, but the chips in this have got to be expensive. I almost wonder if the absolute base model of this machine has much noticeably lower margins than a normal Apple product because that. But they expect/know that most everyone who buys one is going to spec it up.
Re: Apple M3 Ultra
#808Re: Apple M3 Ultra
#809Thunderbolt 5 (TB 5) is pretty handy, you can have a very thin and lightweight laptop, then can get access to external GPU or eGPU via TB 5 if needed [1]. Now you can have your cake (lightweight laptop) and eat it too (potent GPU). [1] Asus just announced the world’s first Thunderbolt 5 eGPU: https://www.theverge.com/24336135/asus-thunderbolt-5-externa...
When I connect to my Mac Studio via Macbook I can select that mode, then change the Displays setting to Dynamic Resolution and then my 'thin client':
- Is fullscreen using the entire 16:10 Macbook screen
- Gets 60 fps low latency performance (including on actual games)
- Transfers audio, I can attend meetings in this mode
- Blanks the host Mac Studio screen
All things that were impossible via VNC - RDP is much better but this new High Performance Screen Share is even more powerful.
The thin lightweight laptop that remotes into a loaded machine has always been my idea of high mobility instead of suffering a laptop running everything locally. This works via LTE as well with some firewall setup.
Re: Apple M3 Ultra
#810Earlier quoted context omitted.
Yep, it's apples to oranges. But sometimes you want apples, and sometimes you want oranges, so it's all good! There's a wide spectrum of potential requirements between memory capacity, memory bandwidth, compute speed, compute complexity, and compute parallelism. In the past, a few GB was adequate for tasks that we assigned to the GPU, you had enough storage bandwidth to load the relevant scene into memory and generat…
Sure, if you want to do training get an NVIDIA card. My point is that it's not worth comparing either Mac or CPU x86 setup to anything with NVIDIA in it. For inference setups, my point is that instead of paying $10000-$15000 for this Mac you could build an x86 system for The "+$4000" for 512GB on the Apple configurator would be "+$1000" outside the Apple world.
That requires an otherwise equivalent PC to exist. I haven’t seen anyone name a PC with a half-TB of unified memory in this thread.
Yeah it’s $4k. Yeah that’s nuts. But it’s the only game in town like that. If the replacement is a $40k setup from Nvidia or whatever that’s a bargain.