Earlier quoted context omitted.
> 96GB of weights. You won't be able to run this on your home GPU. This seems like a non-sequitur. Doesn't MoE select an expert for each token? Presumably, the same expert would frequently be selected for a number of tokens in a row. At that point, you're only running a 7B model, which will easily fit on a GPU. It will be slower when "swapping" experts if you can't fit them all into VRAM at the same time, but it shou…
Someone smarter will probably correct me, but I don’t think that is how MoE works. With MoE, a feed-forward network assesses the tokens and selects the best two of eight experts to generate the next token. The choice of experts can change with each new token. For example, let’s say you have two experts that are really good at answering physics questions. For some of the generation, those two will be selected. But lat…
Mistral "Mixtral" 8x7B 32k model [magnet]
21–30 of 255 posts
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#22Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#23No public statement from Mistral yet. What we know: - Mixture of Experts architecture. - 8x 7B parameters experts (potentially trained starting with their base 7B model?). - 96GB of weights. You won't be able to run this on your home GPU.
> 96GB of weights. You won't be able to run this on your home GPU. This seems like a non-sequitur. Doesn't MoE select an expert for each token? Presumably, the same expert would frequently be selected for a number of tokens in a row. At that point, you're only running a 7B model, which will easily fit on a GPU. It will be slower when "swapping" experts if you can't fit them all into VRAM at the same time, but it shou…
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#24Might be the training code related with the model https://github.com/mistralai/megablocks-public/tree/pstock/m...
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#25Earlier quoted context omitted.
however, if you need to swap experts on each token, you might as well run on cpu.
> Presumably, the same expert would frequently be selected for a number of tokens in a row In other words, assuming you ask a coding question and there's a coding expert in the mix, it would answer it completely.
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#26Earlier quoted context omitted.
> 96GB of weights. You won't be able to run this on your home GPU. This seems like a non-sequitur. Doesn't MoE select an expert for each token? Presumably, the same expert would frequently be selected for a number of tokens in a row. At that point, you're only running a 7B model, which will easily fit on a GPU. It will be slower when "swapping" experts if you can't fit them all into VRAM at the same time, but it shou…
Someone smarter will probably correct me, but I don’t think that is how MoE works. With MoE, a feed-forward network assesses the tokens and selects the best two of eight experts to generate the next token. The choice of experts can change with each new token. For example, let’s say you have two experts that are really good at answering physics questions. For some of the generation, those two will be selected. But lat…
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#27I had to manually add these trackers and now it works: https://gist.github.com/mcandre/eab4166938ed4205bef4
Re: Mistral "Mixtral" 8x7B 32k model [magnet]
#28Some companies spend weeks on landing pages, demos and cute thought through promo videos and then there is Mistral, casually dropping a magnet link on Friday.