Cerebras will be the cheapest way, I think. Let's do the math.
Wonder if crowdfunding could be used to gain shared access to this
131–140 of 157 posts
Cerebras will be the cheapest way, I think. Let's do the math.
Wonder if crowdfunding could be used to gain shared access to this
I've often wondered why a service doesn't exist that allows you to rent out your graphics card for the large data processing needed for training models. Like mining bitcoin except you are doing something actually useful and getting paid actual money for it. Example: - Company Alpha needs $40,000,000 worth of cloud computing for their training model - Company Beta provides them said cloud computing for $30,000,000 fro…
- Fluidstack: https://fluidstack.io
- Vast: https://vast.ai
- QBlocks: https://qblocks.cloud
- RunPod: https://runpod.io
- Sonm: https://sonm.com (blockchain)
- Golem: https://golem.network (blockchain)
- Rentaflop: https://rentaflop.com (rendering specific, blockchain)
- RNDR: https://rendertoken.com (rendering specific, blockchain)
If you want HPC specific cloud providers:
- Crusoe Cloud: https://crusoecloud.com
- Coreweave: https://coreweave.com
- Lambda Labs: https://lambdalabs.com
- Paperspace: https://paperspace.com
As others have pointed out, the decentralized clouds can't offer high performance interconnects (e.g. InfiniBand) that a lot of folks are using for LLM training. There are definitely initiatives underway to reduce dependence on these interconnects and build performant distributed training (again, some threads below mention this), but I think it's mostly academic at this point.
Disclosure: I run product at Crusoe Cloud, which aims to provide ML training at half the cost of a hyperscaler, while also being carbon reducing (https://crusoecloud.com/climate-impact/).
It's not that easy. Access to enough compute is one thing. However, you also need a proper dataset (beyond Common Crawl and Wikipedia), excellent research expertise and engineering capabilities. So even if you throw money or free credits for cloud compute out there it will not be enough. We've seen this happen with EleutherAI who were not capable of reaching their initial target of "replicating" GPT-3 and could only…
We solved the proper dataset part at least. https://arxiv.org/abs/2101.00027 My contribution was around 19,000 books.
what does this mean? not meaning to cross examine you, just curious how people contribute to The Pile since it seemingly appeared out of nowhere
It's not that easy. Access to enough compute is one thing. However, you also need a proper dataset (beyond Common Crawl and Wikipedia), excellent research expertise and engineering capabilities. So even if you throw money or free credits for cloud compute out there it will not be enough. We've seen this happen with EleutherAI who were not capable of reaching their initial target of "replicating" GPT-3 and could only…
is there a crowdsourced list of text corpuses somewhere? i bet thats the starting point for all this. i'm only aware of C4 and The Pile.
I've often wondered why a service doesn't exist that allows you to rent out your graphics card for the large data processing needed for training models. Like mining bitcoin except you are doing something actually useful and getting paid actual money for it. Example: - Company Alpha needs $40,000,000 worth of cloud computing for their training model - Company Beta provides them said cloud computing for $30,000,000 fro…
If there was a feasible crowdfunded solution to this and putting it in the hands of the people - I would certainly be prepared to lay down up to £5K.
Earlier quoted context omitted.
In addition to privacy, performance, and portability, also: * Servers in a datacenter are much more reliable than a network of PCs (power goes off, someone decides to play Crysis, etc) * People will find ways to scam you (pretend like they’re doing the calculation while not actually doing it) * Economies of scale means a datacenter will probably be cheaper than what you’d have to pay the PC owners (power consumption,…
What if you remove the financial incentive? I'd contribute my GPU time to a Folding@Home style project if it meant that we had powerful, open LLMs that were free to use. I'm positive many others would as well. As far as worrying about scammers, could you send the same compute task and training data to multiple clients and validate the results against each other? If they differed, you could throw the results and try a…
By sending the work to 3 people each time, you’re effectively cutting your (already limited) resource pool by 66%.
Earlier quoted context omitted.
We solved the proper dataset part at least. https://arxiv.org/abs/2101.00027 My contribution was around 19,000 books.
> My contribution was around 19,000 books what does this mean? not meaning to cross examine you, just curious how people contribute to The Pile since it seemingly appeared out of nowhere
I was convinced that a model needed to be able to read like we do. And what do we do when we read? Pick up a book.
That turns out to be surprisingly hard, at least for training data. Step one is to acquire the books. Step two is to turn them into a readable format for computers.
Both steps were very hard. I lucked out on step one because The Eye happened to host all of bibliotok, which came to around 30k books or so.
Trouble is, lots of those are PDFs. And although humans are great at reading those, they fucking suck for blind people. And a gpt is a blind person in a sense, because it needs to follow a linear sequence of words — something that PDFs are horrible at giving.
But one day I realized that epubs were merely html files, and aaronsw happened to write an amazing html to text converter. I had to hack it to fix a few corner cases. But after a few days, I ran it across all 19,000 epubs I spidered, then zipped the whole thing up and called it books3: https://twitter.com/theshawwn/status/1320282149329784833?s=4...
It’s one of the larger components of the pile, I think around 35%. Which is quite the hefty sum when it’s purely text. I still have a hard time wrapping my head around just how mindbogglingly big 800GB of text is.
Earlier quoted context omitted.
We solved the proper dataset part at least. https://arxiv.org/abs/2101.00027 My contribution was around 19,000 books.
isn't Common Crawl much, much larger than this? ~6 pebibytes from what I remember
bmk was the magician there. https://twitter.com/nabla_theta?s=21&t=Gt6YrATJHnmY046MdzhYD...
Earlier quoted context omitted.
> My contribution was around 19,000 books what does this mean? not meaning to cross examine you, just curious how people contribute to The Pile since it seemingly appeared out of nowhere
Not at all, I love talking about it. I was convinced that a model needed to be able to read like we do. And what do we do when we read? Pick up a book. That turns out to be surprisingly hard, at least for training data. Step one is to acquire the books. Step two is to turn them into a readable format for computers. Both steps were very hard. I lucked out on step one because The Eye happened to host all of bibliotok,…
this maaaay be covered in the Pile's writeup (which i have not yet read) but i wonder who was curating the overall "mix" of the content. seems easily biased to, say, public domain books, since the corpus is easily available.
when people say things like "GPT3 has been trained on all of the internet" i suspect this is a gross exaggeration. In reality it's just C4/commoncrawl, so that's like 800GB of text.