Live data from Hacker News

Cerebras-GPT: Open Compute-Optimal Language Models Trained on Cerebras Cluster

arxiv.org

11–13 of 13 posts

Re: Cerebras-GPT: Open Compute-Optimal Language Models Trained on Cerebras Cluster

#11

Recently, we announced in this post ( https://news.ycombinator.com/item?id=35343763#35345980 ) the release of Cerebras-GPT — a family of open-source GPT models trained on the Pile dataset using the Chinchilla formula. Today, we are excited to announce the availability of the Cerebras-GPT research paper on arXiv.

Thanks for publishing this. I quickly skimmed the paper, I saw the impressive linear scaling as you scaled to 16 nodes. How long did it take to train the various models in wall clock time?

Re: Cerebras-GPT: Open Compute-Optimal Language Models Trained on Cerebras Cluster

#13
Very interesting that someone finally tries out muP in the real world. Do I understand the usage correctly:

MuP is only used to get around choosing an lr for each size? Here I wonder how it compares to standard heuristics like the one in the OG scaling laws paper by OAI and tricks like back winding a few steps after loss explosion.

For some reason muP was not trusted with the largest trainings? Why is that?

Post reply on HN