Earlier quoted context omitted.
Which means this is a whole new chip. It may be M3 based, but with added interposer support and new thunderbolt stuff. Which, at this point, why not just use M4 as a base?
Could be that M4 requires a different TSMC fab that is at full production doing iPhones.
Apple M3 Ultra
631–640 of 1001 posts
Re: Apple M3 Ultra
#632Whoa. M3 instead of M4. I wonder if this was basically binning, but I thought that I had read somewhere that the interposer that enabled this for the M1 chips where not available. That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use.…
> This hardware is really being held back by the operating system at this point. Apple could either create a 2U rack hardware and support Linux (and I mean Apple supporting it, not hobbysts), or have a build of Darwin headless that could run on that hardware. But in the later case, we probably wouldn't have much software available (though I am sure people would eventually starting porting software to it, there is alr…
Re: Apple M3 Ultra
#633512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions. I'm curious about how…
They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand. The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new…
If they didn't increase the memory bandwidth, then 512GB will enable longer context lengths and that's about it right? No speedups
For any speedups You may need some new variant of FlashAttention3 or something along similar lines to be purpose built for Apple GPUs.
Re: Apple M3 Ultra
#634Earlier quoted context omitted.
Yep, it's apples to oranges. But sometimes you want apples, and sometimes you want oranges, so it's all good! There's a wide spectrum of potential requirements between memory capacity, memory bandwidth, compute speed, compute complexity, and compute parallelism. In the past, a few GB was adequate for tasks that we assigned to the GPU, you had enough storage bandwidth to load the relevant scene into memory and generat…
> we still work with powers of two. Please. We do. Common people don't. It's easier to write "over half a terabyte" than explain (again) to millions of people what the power of two is.
Re: Apple M3 Ultra
#635Earlier quoted context omitted.
> we still work with powers of two. Please. We do. Common people don't. It's easier to write "over half a terabyte" than explain (again) to millions of people what the power of two is.
Anyone who calls 512 gigs "over half a terabyte" is bullshitting. No, thank you.
Re: Apple M3 Ultra
#636512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions. I'm curious about how…
"unified memory" funny that people think this is so new, when CRAY had Global Heap eons ago...
Re: Apple M3 Ultra
#637Earlier quoted context omitted.
> that 10k is a absolute bargain The higher end NVidia workstation boxes won’t run well on normal 20amp plugs. So you need to move them to a computer room (whoops, ripped those out already) or spend months getting dedicated circuits run to office spaces.
In the US, normal circuits aren't always 20A, especially in residential buildings, where they are more commonly 15A in bedrooms and offices. https://en.wikipedia.org/wiki/NEMA_connector
Re: Apple M3 Ultra
#638Earlier quoted context omitted.
How would you compare the tok/sec between this setup and the M3 Max?
3.5 - 4.5 tokens/s on the $2,000 AMD Epyc setup. Deepseek 671b q4. The AMD Epyc build is severely bandwidth and compute constrained. ~40 tokens/s on M3 Ultra 512GB by my calculation.
If the M3 can run 24/7 without overheating it's a great deal to run agents. Especially considering that it should run only using 350W... so roughly $50/mo in electricity costs.
Re: Apple M3 Ultra
#639Earlier quoted context omitted.
Haven't the Max/Ultra type chips always come much later, close to when the next number of standard chips came out? M2 Max was not available when M2 launched, for example.
An Ultra has never come out after the next gen base model, let alone the next gen Pro/Max model before. M1: November 10, 2020 M1 Pro: October 18, 2021 M1 Max: October 18, 2021 M1 Ultra: March 8, 2022 ------------------------- M2: June 6, 2022 M2 Pro: January 17, 2023 M2 Max: January 17, 2023 M2 Ultra: June 5, 2023 ------------------------- M3: October 30, 2023 M3 Pro: October 30, 2023 M3 Max: October 30, 2023 -------…