Live data from Hacker News

Apple M3 Ultra

apple.com

581–590 of 1001 posts

Re: Apple M3 Ultra

#581

Earlier quoted context omitted.

Also should note that 800/819GB/s of memory bandwidth is actually VERY usable for LLMs. Consider that a 4090 is just a hair above 1000GB/s

Does it work like that though at this larger scale? 512GB of VRAM would be across multiple NVIDIA cards, so the bandwidth and access is parallelized. But here it looks more of a bottleneck from my (admittedly naive) understanding.

For inference the bandwidth is generally not parallelized because the weights need to go through the model layer by layer. The most common model splitting method is done by assigning each GPU a subset of the LLM layers and it doesn't take much bandwidth to send model weights via PCIE to the next GPU.

Re: Apple M3 Ultra

#583
post #47

Earlier quoted context omitted.

>I had read somewhere that the interposer that enabled this for the M1 chips where not available. With all my love and respect for "Apple rumors" writers; this was always "I read five blogposts about CPU design and now I'm an expert!" territory. The speculation was based on the M3 Maxes die shots not having the interposer visible, which... implies basically nothing whether that _could have_ been supported in an M3 Ul…

I’m guessing it’s not really a M3. No M3 has thunderbolt 5. This is a new chip with M3 marketing. I’d expect this from Intel, not Apple.

TB 5 seems like the sort of thing you could 'slap on' to a beefy enough chip.

Or the sort of thing you put onto a successor when you had your fingers crossed that the spec and hardware would finalize in time for your product launch but the fucking committee went into paralysis again at the last moment and now your product has to ship 4 months before you can put TB 5 hardware on shelves. So you put your TB4 circuitry on a chip that has the bandwidth to handle TB5 and you wait for the sequel.

Re: Apple M3 Ultra

#584
post #564

Two questions for the fellow HNers: 1. What are various average joe (as opposed to researchers, etc.) use cases for running powerful AI models locally vs. just using cloud AI. Privacy of course is a benefit, but it by itself may not justify upgrades for an average user. Or are we expecting that new innovation will lead to much more proliferation of AI and use cases that will make running locally more feasible? 2. Wit…

I don’t currently use AIs, but if I did, they would be local. Simply put: I can’t build my professional career around tools that I do not own.

Re: Apple M3 Ultra

#585

Whoa. M3 instead of M4. I wonder if this was basically binning, but I thought that I had read somewhere that the interposer that enabled this for the M1 chips where not available. That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use.…

> This hardware is really being held back by the operating system at this point. It really is. Even if they themselves won't bring back their old XServe OS variant, I'd really appreciate it if they at least partnered with a Linux or BSD (good callout, ryao) dev to bring a server OS to the hardware stack. The consumer OS, while still better (to my subjective tastes) than Windows, is increasingly hampered by bloat and…

I miss the XServe almost as much as I miss the Airport Extreme.

Re: Apple M3 Ultra

#586
post #241

Earlier quoted context omitted.

For enterprise markets, this is table stakes. A lot of datacenter customers will probably ignore this release altogether since there isn't a high-bandwidth option for systems interconnect.

The Mac Studio isn’t meant for data centers anyway? It’s a small and silent desktop form factor — in every respect the opposite of a design you’d want to put in a rack. A long time ago Apple had a rackmount server called Xserve, but there’s no sign that they’re interested in updating that for the AI age.

Apple recently announced they’re building a new plant in Texas to produce servers. Yes, they need servers for their Private Compute Cloud used by Apple Intelligence, but it doesn’t only need to be for that.

From https://www.apple.com/newsroom/2025/02/apple-will-spend-more...

As part of its new U.S. investments, Apple will work with manufacturing partners to begin production of servers in Houston later this year. A 250,000-square-foot server manufacturing facility, slated to open in 2026, will create thousands of jobs.

Re: Apple M3 Ultra

#587
post #342

Earlier quoted context omitted.

Other than the NPU, it’s not really a game changer; here’s a 512GB AMD deepseek build for $2000: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...

between 4.25 to 3.5 TPS (tokens per second) on the Q4 671b full model. 3.5 - 4.25 tokens/s. You're torturing yourself. Especially with a reasoning model. This will run it at 40 tokens/s based on rough calculation. Q4 quant. 37b active parameters. 5x higher price for 10x higher performance.

Also you don't have to deal with Windows. Which people who do not understand Apple are very skilled at not noticing.

If you've ever used git, svn, or an IDE side by side on corporate Windows versus Apple I don't know why you would ever go back.

Re: Apple M3 Ultra

#588
post #564

Two questions for the fellow HNers: 1. What are various average joe (as opposed to researchers, etc.) use cases for running powerful AI models locally vs. just using cloud AI. Privacy of course is a benefit, but it by itself may not justify upgrades for an average user. Or are we expecting that new innovation will lead to much more proliferation of AI and use cases that will make running locally more feasible? 2. Wit…

IMO it's all about privacy. Perhaps also availability if the main LLM providers start pulling shenanigans but it seems like that's not going to be a huge problem with how many big players are in the space. I think a great use case for this would be in a company that doesn't want all of their employees sending LLM queries about what they're working on outside the company. Buy one or two of these and give everybody a c…

I’ll add to this that while I couldn’t care less about open AI seeing my general coding questions, I wouldn’t run actual important data through ChatGPT.

With a local model, I could toss anything in there. Database query outputs, private keys, stuff like that. This’ll probably become more relevant as we give LLM’s broader use over certain systems.

Like right now I still mostly just type or paste stuff into ChatGPT. But what about when I have a little database copilot that needs to read query results, and maybe even run its own subset of queries like schema checks? Or some open source computer-use type thingy needs to click around in all sorts of places I don’t want openAI going, like my .env or my bash profile? That’s the kinda thing I’d only use a local model for

Re: Apple M3 Ultra

#589
post #549

$14K with 512gb memory and 16 Tb storage

I cannot believe I’m saying that, but: for apple that’s rather cheap. Threadripper boxes with that amount of memory do not come a lot cheaper. Considering what apples pricing when it comes to memory in other devices, 4K for the 96GB to 512GB upgrade is a bargain.

Re: Apple M3 Ultra

#590
post #169

Earlier quoted context omitted.

NVIDIA RTX 4090: ~1,008 GB/s NVIDIA RTX 4080: ~717 GB/s AMD Radeon RX 7900 XTX: ~960 GB/s AMD Radeon RX 7900 XT: ~800 GB/s How's that slow exactly ? You can have 10000000Gb/s and without enough VRAM it's useless.

I have a 4090 and, out of curiosity, I looked up the FLOPS in comparison with Apple chips. Nvidia RTX 4090 (Ada Lovelace) FP32: Approximately 82.6 TFLOPS FP16: When using its 4th‑generation Tensor Cores in FP16 mode with FP32 accumulation, it can deliver roughly 165.2 TFLOPS (in non‑tensor mode, the FP16 rate is similar to FP32). FP8: The Ada architecture introduces support for an FP8 format; using this mode (again w…

> Anything that knocks Nvidia down a notch is good for humanity.

I don't love Nvidia a whole lot but I can't understand where this sentinent comes from. Apple abandoned their partnership with Nvidia, tried to support their own CUDA alternative with blackjack and hookers (OpenCL), abandoned that, and began rolling out a proprietary replacement.

CUDA sucks for the average Joe, but Apple abandoned any chance of taking the high road when they cut ties with Khronos. Apple doesn't want better AI infrastructure for humanity; they envy the control Nvidia wields and want it for themselves. Metal versus CUDA is the type of competition where no matter who wins, humanity loses. Bring back OpenCL, then we'll talk about net positives again.

Post reply on HN