Whoa. M3 instead of M4. I wonder if this was basically binning, but I thought that I had read somewhere that the interposer that enabled this for the M1 chips where not available. That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use.…
Other than the NPU, it’s not really a game changer; here’s a 512GB AMD deepseek build for $2000: https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...
between 4.25 to 3.5 TPS (tokens per second) on the Q4 671b full model.
3.5 - 4.25 tokens/s. You're torturing yourself. Especially with a reasoning model.This will run it at 40 tokens/s based on rough calculation. Q4 quant. 37b active parameters.
5x higher price for 10x higher performance.