Live data from Hacker News

Apple M3 Ultra

apple.com

421–430 of 1001 posts

Re: Apple M3 Ultra

#421
post #342

Whoa. M3 instead of M4. I wonder if this was basically binning, but I thought that I had read somewhere that the interposer that enabled this for the M1 chips where not available. That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use.…

Other than the NPU, it’s not really a game changer; here’s a 512GB AMD deepseek build for $2000: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...

  between 4.25 to 3.5 TPS (tokens per second) on the Q4 671b full model.
3.5 - 4.25 tokens/s. You're torturing yourself. Especially with a reasoning model.

This will run it at 40 tokens/s based on rough calculation. Q4 quant. 37b active parameters.

5x higher price for 10x higher performance.

Re: Apple M3 Ultra

#422
post #393

Earlier quoted context omitted.

They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand. The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new…

> The the question is if a llm will run with usable performance at that scale? This is the big question to have answered. Many people claim Apple can now reliably be used as a ML workstation, but from the numbers I've seen from benchmarks, the models may fit in memory, but the performance for tok/sec is so slow to not feel worth it, compared to running it on NVIDIA hardware. Although it be expensive as hell to get 51…

It is much slower than nVidia, but for a lot of personal-use LLM scenarios, it's very workable. And it doesn't need to be anywhere near as fast considering it's really the only viable (affordable) option for private, local inference, besides building a server like this, which is no faster: https://news.ycombinator.com/item?id=42897205

Re: Apple M3 Ultra

#423
post #272

Earlier quoted context omitted.

Nah, if I ever wrote an article about the software crisis on the Linux desktop, there’d be flames here making Apple’s issues look small.

It'd be an interesting flame war in the comments, if nothing else, go for it! I'm happy to give plenty of concrete evidence why Linux is more suitable for professionals than macOS is in 2025 :)

Try copy pasting bash snippets between any linux text editor and terminal.

Now try the same with notes on a mac. Notes mangles the punctuation and zsh is not bash.

Re: Apple M3 Ultra

#424
post #152

apple keeps talking about the Neural Engine. Does anything actually use it? Seems like all the current LLM and Stable Diffusion packages (including MLX) use the GPU.

Yeah I agree.

The Neural Engine is useful for a bunch of Apple features, but seems weirdly useless for any LLM stuff... been wondering if they'd address it on any of these upcoming products. AI is so hype right now it seems odd that they have specialised processor that doesn't get used for the kind of AI people are doing. I can see in the latest release:

> Mac Studio is a powerhouse for AI, capable of running large language models (LLMs) with over 600 billion parameters entirely in memory, thanks to its advanced GPU

https://www.apple.com/newsroom/2025/03/apple-unveils-new-mac...

i.e. LLMs still run on the GPU not the NPU

Re: Apple M3 Ultra

#426
post #325

Earlier quoted context omitted.

They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand. The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new…

Guess what? I'm on a mission to completely max out all 512GB of mem...maybe by running DeepSeek on it. Pure greed!

You could always just open a few Chrome tabs…

Re: Apple M3 Ultra

#427

Earlier quoted context omitted.

If Apple supported Linux (headless) natively, and we could rack m4 pros, I absolutely would use them in our Colo. The CPUs have zero competition in terms of speed, memory bandwidth. Still blown away no other company has been able to produce Arm server chips that can compete.

If I read this right, the r8g.48xlarge at AMZN [1] has 192 cores and 1536GB which exceeds the M3 Ultra in some metrics. It reminds me of the 1990s when my old school was using Sun machines based on the 68k series and later SPARC and we were blown away with the toaster-sized HP PA RISC machine that was used for student work for all the CS classes. Then Linux came out and it was clear the 386 trashed them all in terms…

It seems Graviton 4 CPUs have 12-channels of DDR5-5600 i.e 540GB/s main memory bandwidth for the CPU to use. M3 Ultra has 64-channels of LPDDR5-6400 i.e. ~800GB/s of memory bandwidth for the CPU or the GPU to use. So the M3 Ultra has way fewer (CPU) cores, but way more memory bandwidth. Depends what you're doing.

Re: Apple M3 Ultra

#428
post #280

I am confused. I got an M4 with 64 GB Ram. Did I buy something from the future? :) Now why M3 now? Not M4 Ultra.

Haven't the Max/Ultra type chips always come much later, close to when the next number of standard chips came out? M2 Max was not available when M2 launched, for example.

An Ultra has never come out after the next gen base model, let alone the next gen Pro/Max model before.

M1: November 10, 2020

M1 Pro: October 18, 2021

M1 Max: October 18, 2021

M1 Ultra: March 8, 2022

-------------------------

M2: June 6, 2022

M2 Pro: January 17, 2023

M2 Max: January 17, 2023

M2 Ultra: June 5, 2023

-------------------------

M3: October 30, 2023

M3 Pro: October 30, 2023

M3 Max: October 30, 2023

-------------------------

M4: May 7, 2024

M4 Pro: October 30, 2024

M4 Max: October 30, 2024

-------------------------

M3 Ultra: March 5, 2025

Re: Apple M3 Ultra

#429

Earlier quoted context omitted.

> lack of robust python support There is no such thing. Tell me, which combination of the 15+ virtual environments, dependency management and Python version managers would you use? And how would you prevent "project collision" (where one Python project bumps into another one and one just stops working)? Example: SSL library differences across projects is a notorious culprit. Python is garbage and I don't understand w…

Virtualenv’s been a thing for many years, it’s built into Python, and it solves all that without adding additional tooling. And if you’re genuinely asking, everything’s converging toward uv. If you pick only one, use that and be done with it.

I’ve been using virtualenv for a decade, and we use uv at work.

Neither fixed anything. They just make it slightly less painful to deal with python scripts’ constant bitrot.

They also make python uniquely difficult to dockerize.

Re: Apple M3 Ultra

#430
post #307

Earlier quoted context omitted.

The ANE and tensor cores are not comparable though. One is literally meant for low cost inference while the others are meant for acceleration of training. If you squint, yeah they look the same, but so does the microcontroller on the GPU and a full blown CPU. They’re fundamentally different purposes, architectures and scale of use. The ANE can’t even really be used directly. Apple heavily restricts the use via CoreML…

So now the TOPS are not comparable because M3 is much slower than an Nvidia GPU? That's not how comparisons work. My numbers are correct, the M3 Ultra has around 1 % of the TOPS performance of a RTX 5090. Comparing against the GPU would look even worse for apple. Do you think Apple added the neural engine just for fun? This is exactly what the neural engine is there for.

You’re completely missing the point. The ANE is not equivalent as a component to the tensor cores. It has nothing to do with comparison of TOPs but as what they’re intended for.

Try and use the ANE in the same way you would use the tensor cores. Hint: you can’t, because the hardware and software will actively block you.

They’re meant for fundamentally different use cases and power loads. Even apples own ML frameworks do not use the ANE for anything except inference.

Post reply on HN