Earlier quoted context omitted.
OK, I am curious now: What kind of hardware would I need to run such a model for a couple of users with decent performance? Where could I get a mapping of token / time vs hardware?
Unsure if anyone has specific hardware benchmarks for the 405b model yet, since it's so new, but elsewhere in this thread I outlined a build that'd probably be capable of running a quantized version of Llama 3.1 405b for roughly $10k. The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU…
I would be curious to see relative failure rates over time of consumer vs Quadro cards as well.