Implementation of mixture of experts language model in a single file of PyTorch
1–10 of 15 posts
Re: Implementation of mixture of experts language model in a single file of PyTorch
#2Re: Implementation of mixture of experts language model in a single file of PyTorch
#3A from scratch implementation of a sparse mixture of experts language model in a single file of PyTorch. This is inspired by and largely based on Andrej Karpathy's project 'makemore' and borrows a number of re-usable components from that implementation. Just like makemore, makeMoE is also an autoregressive character-level language model but uses the aforementioned sparse mixture of experts architecture. I added Exper…
Adding scaled unit gaussian noise to the logits
noise = torch.randn_like(logits)*F.softplus(noise_logits)
noisy_logits = logits + noise
Question, if you changed this Gaussian normal for Gumbel noise you would get something like Gumbel softmax, right? I'm curious why not use it? Isn't it a usual way to implement differentiable discrete selection? My curiosity is about the effectiveness of Gumbel softmax since I have had some trouble using it in practice so I'm curious why it's not used here and if there are downsides to it compared to other methods. Honestly just adding normal noise like this seems simpler anyway.Re: Implementation of mixture of experts language model in a single file of PyTorch
#4Re: Implementation of mixture of experts language model in a single file of PyTorch
#5Very cool. I'm curious - did you find the results from your mixture of experts model to be (qualitatively) better than with the standard approach?
Re: Implementation of mixture of experts language model in a single file of PyTorch
#6A from scratch implementation of a sparse mixture of experts language model in a single file of PyTorch. This is inspired by and largely based on Andrej Karpathy's project 'makemore' and borrows a number of re-usable components from that implementation. Just like makemore, makeMoE is also an autoregressive character-level language model but uses the aforementioned sparse mixture of experts architecture. I added Exper…
Adding scaled unit gaussian noise to the logits noise = torch.randn_like(logits)*F.softplus(noise_logits) noisy_logits = logits + noise Question, if you changed this Gaussian normal for Gumbel noise you would get something like Gumbel softmax, right? I'm curious why not use it? Isn't it a usual way to implement differentiable discrete selection? My curiosity is about the effectiveness of Gumbel softmax since I have h…
Re: Implementation of mixture of experts language model in a single file of PyTorch
#7Re: Implementation of mixture of experts language model in a single file of PyTorch
#8Similar MoE implementation was on GitHub for a while, since Jan 2024 https://github.com/zxaall/moegpt
Re: Implementation of mixture of experts language model in a single file of PyTorch
#9Similar MoE implementation was on GitHub for a while, since Jan 2024 https://github.com/zxaall/moegpt
Oh nice. What's new here would be noisy top-k routing and expert capacity. It also seems to use the nanoGPT base from Andrej Karpathy. Mine is from January as well. Here's the original blog: https://huggingface.co/blog/AviSoori1x/makemoe-from-scratch
Re: Implementation of mixture of experts language model in a single file of PyTorch
#10Very cool. I'm curious - did you find the results from your mixture of experts model to be (qualitatively) better than with the standard approach?