Earlier quoted context omitted.
I cannot express how dirt cheap that pricepoint is for what's on offer, especially when you're comparing it to rackmount servers. By the time you've shoehorned in an nVidia GPU and all that RAM, you're easily looking at 5x that MSRP; sure, you get proper redundancy and extendable storage for that added cost, but now you also need redundant UPSes and have local storage to manage instead of centralized SANs or NASes. F…
> By the time you've shoehorned in an nVidia GPU and all that RAM, you're easily looking at 5x that MSRP That nvidia GPU setup will actually have the compute grunt to make use of the RAM, though, which this M3 Ultra probably realistically doesn't. After all, if the only thing that mattered was RAM then the 2TB you can shove into an Epyc or Xeon would already be dominating the AI industry. But they aren't, because it…
Write me an AWS CloudFormation file that does the following:
* Deploys an Amazon Kubernetes Cluster
* Deploys Busybox in the namespace "Test1", including creating that Namespace
* Deploys a second Busybox in the namespace "Test3", including creating that Namespace
* Creates a PVC for 60GB of storage
The M1Pro laptop with 16GB of Unified Memory: * 21.28 seconds for "Thinking"
* 0.22s to the first token
* 18.65 tokens/second over 1484 tokens in its responses
* 1m:23s from sending the input to completion of the output
The 10900k CPU, with 64GB of RAM and a full-fat RTX 3090 GPU in it: * 10.88 seconds for "thinking"
* 0.04s to first token
* 58.02 tokens/second over 1905 tokens in its responses
* 0m:34s from sending the input to completion of the output
Same model, same loader, different architectures and resources. This is why a lot of the AI crowd are on Macs: their chip designs, especially the Neural Engine and GPUs, allow quite competent edge inference while sipping comparative thimbles of energy. It's why if I were all-in on LLMs or leveraged them for work more often (which I intend to, given how I'm currently selling my generalist expertise to potential employers), I'd be seriously eyeballing these little Mac Studios for their local inference capabilities.