Live data from Hacker News

Mistral "Mixtral" 8x7B 32k model [magnet]

twitter.com

41–50 of 255 posts

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#41
post #6

No public statement from Mistral yet. What we know: - Mixture of Experts architecture. - 8x 7B parameters experts (potentially trained starting with their base 7B model?). - 96GB of weights. You won't be able to run this on your home GPU.

>> You won't be able to run this on your home GPU.

Would this allow you to run each expert on a cheap commodity GPU card so that instead of using expensive 200GB cards we can use a computer with 8 cheap gaming cards in it?

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#42
post #28

Some companies spend weeks on landing pages, demos and cute thought through promo videos and then there is Mistral, casually dropping a magnet link on Friday.

I'm curious about their business model.

They can make plenty by offering consulting fees for finetuning and general support around their models.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#43
post #3

Still 7B, but now with 32k context. Looking forward to see how it compares with the previous one, and what the community does with it.

Not 7B, 8x7B. It will run with the speed of a 7B model while being much smarter but requiring ~24GB of RAM instead of ~4GB (in 4bit).

Given the config parametes posted, its 2 experts per token, so the conputation cost per token should be the cost of the conponent that selects experts + 2× cost of a 7B model.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#45
post #37

Mistral sure does not bother too much with explanations, but this style gives me much more confidence in the product than Google's polished, corporate, soulless announcement of Gemini!

I will take weights over docs.

Its does remind me how some Google employee was bragging that they disclosed the weights for the Gemini, and only the small mobile Gemini, as if that's a generous step over other companies.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#47
post #42
post #28

Earlier quoted context omitted.

I'm curious about their business model.

They can make plenty by offering consulting fees for finetuning and general support around their models.

"plenty" is not a word some of these people understand however

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#48

Some companies spend weeks on landing pages, demos and cute thought through promo videos and then there is Mistral, casually dropping a magnet link on Friday.

I'm sure it's also a marketing move to build a certain reputation. Looks like it's working.

Not geoblocking the entirety of Europe also makes them stand out like a ringmaster amongst clowns.

Re: Mistral "Mixtral" 8x7B 32k model [magnet]

#50
post #6

No public statement from Mistral yet. What we know: - Mixture of Experts architecture. - 8x 7B parameters experts (potentially trained starting with their base 7B model?). - 96GB of weights. You won't be able to run this on your home GPU.

at 4 bits you could run it on a 3090 right?
Post reply on HN