TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
1–10 of 61 posts
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#2Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#3They are training the model on 3000/22=136 times the value of the chinchilla scale. It will be interesting to see how much it will improve after way beyond this value.
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#4>It means you can train a chinchilla-optimal TinyLlama (1.1B param, 22B tokens) in 32 hours with 8 A100. They are training the model on 3000/22=136 times the value of the chinchilla scale. It will be interesting to see how much it will improve after way beyond this value.
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#5>It means you can train a chinchilla-optimal TinyLlama (1.1B param, 22B tokens) in 32 hours with 8 A100. They are training the model on 3000/22=136 times the value of the chinchilla scale. It will be interesting to see how much it will improve after way beyond this value.
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#6>It means you can train a chinchilla-optimal TinyLlama (1.1B param, 22B tokens) in 32 hours with 8 A100. They are training the model on 3000/22=136 times the value of the chinchilla scale. It will be interesting to see how much it will improve after way beyond this value.
I now come to understand that the technobable in Star Trek wasn't that well predicted, in the future we will not be reversing polarities by alligning field cores. Picard will have us align our llamas with chiwawas to get an alpacafied chinchilla model.
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#7The link that says you can watch cross-entropy loss live is locked or broken.
Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#8Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#9Re: TinyLlama project aims to pretrain a 1.1B Llama model on 3T tokens
#10What does “pretrain” mean in this context? It sounds like normal training