Earlier quoted context omitted.
Thats a good point. For example both the RTX 4090 and the RTX 6000 Ada Generation use the AD102 chip. The RTX 6000 Ada though, would be able to run 70b models due to the larger memory pool despite having the same memory interface width.
Would the not-yet-released M3 Mac Mini with upgraded RAM be enough to do some beginner level LLM work?
Those supposedly will go up to 24Gb RAM, which is only enough for inference for ~34b quantized models. At which point you might as well get an RTX 3090 for a PC that you probably already have, and get a better bang for the buck.
It also really depends on what you consider "beginner level". Fine-tuning is really easy these days and many people do it, but you really want CUDA for that.