I wish these releases had short but meaningful descriptions for their models.
E.g. 10 layers x 2048 tokens x 1024 embedding model using full attention.
141–143 of 143 posts
E.g. 10 layers x 2048 tokens x 1024 embedding model using full attention.
Earlier quoted context omitted.
Some people use second hand P40 GPUs, which go for around 200-300$. Combine 3 of them with SLI and you've got 72GB of VRAM for less then $1000
I do use a P40 for my machine learning box, but I'm curious how you put three on the same system, given they need a CPU power plug and a pci-e port. Then, to cool them, you need to plug your own cooling system, requiring more specific power plugs to be available. What kind of chassis, motherboard, power unit you use to do that? It'll certainly will cost more than $1000 anyway, especially since you also need a decent…