Live data from Hacker News

Llama 3.1

llama.meta.com

11–20 of 279 posts

Re: Llama 3.1

#13
post #3

Nice, someone donate me a few 4090s :(

Your going to need a lot more than a few, 800G VRAM needed

If previous quantization results hold up, fp8 will have nearly identical performance while using 405GiB for weights, but the KV cache size will still be significant.

Too bad, too, I don't think my PC will fit 20 4090s (480GiB).

Post reply on HN