Earlier quoted context omitted.
Check out the BLOOM models if you want to see a first stab at that. If you can find an economical way of running it though, let me know.
A cluster of 6 year old 24GB NVIDIA Teslas should do the trick...they run for about $100 apiece. Put 12 or so of them together and you have the VRAM for a GPT3 clone.
Amazon has them listed at $200, but still, that's only $2,400 for 12 of them.
Still, adds up once you get the hardware you'd need to NVlink 12 of them, and then on top of that, the price of power/perf you get probably isn't great compared to modern compute.
Wonder what your volume would have to be before getting a box with 8 A100's from Lambdalabs would be the better tradeoff.