I run a lot of local models (I am always experimenting) on my 32G M2-Pro MacMini - I would love to upgrade. The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig. Hybrid ‘small local’…
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
Apple introduces M6 and M5 Ultra
521–530 of 1001 posts
Re: Apple introduces M6 and M5 Ultra
#522I run a lot of local models (I am always experimenting) on my 32G M2-Pro MacMini - I would love to upgrade. The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig. Hybrid ‘small local’…
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
A PC with similar capabilities is going to sound like a jet taking off.
Re: Apple introduces M6 and M5 Ultra
#523How many Neural Accelerators does M5 Ultra has in its 64 vs 80-GPU variants? Anybody knows? Just to clarify, Neural Accelerators have nothing to do with Neural Engine.
Re: Apple introduces M6 and M5 Ultra
#524Earlier quoted context omitted.
Reading Infinite Games, they tell an anecdote about a Microsoft exec on a flight telling an Apple exec that the Zune was a way better portable music player than the iPod. The Apple exec was just “yup, you’re probably right” and then soon after the iPhone dropped
To be fair, the Zune actually was a better media player than the ipod in almost every way. That wasn't a false statement. The issue was the abysmal marketing and "me too" attitude Microsoft had (and still has).
Re: Apple introduces M6 and M5 Ultra
#525I've briefly tested M5 Pro in an Apple store and was surprised by how quick it felt and did anything. A tangible and significant difference, and I really feel it would be good getting it, or M6.
However - before, MacOS was a big factor driving me towards Apple - now, it would be hard for me to give up my Linux setup and all the stuff I love about it, even for such performance...Linux progressed really nicely, while MacOS deteriorated at the same rate, and it's changing the balance and the decision for me.
Re: Apple introduces M6 and M5 Ultra
#526Re: Apple introduces M6 and M5 Ultra
#527Re: Apple introduces M6 and M5 Ultra
#528I run a lot of local models (I am always experimenting) on my 32G M2-Pro MacMini - I would love to upgrade. The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig. Hybrid ‘small local’…
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
Re: Apple introduces M6 and M5 Ultra
#529Earlier quoted context omitted.
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VR…
A Mac mini running an LLM is quiet A PC with similar capabilities is going to sound like a jet taking off.
Not at all. Airflow with big fans is quiet. What I do hear is coil-whine. In fact my PC is quieter than my Macbook when both are running top speed. But one has 4000 AI TOPS.
Re: Apple introduces M6 and M5 Ultra
#530Earlier quoted context omitted.
when was the turing test beaten?
It depends who takes the test. I am not yet, to my knowledge, fooled by AI. I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this? I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved. But in a more practical sense, if AI can impersonat…