Live data from Hacker News

Mistral AI Launches New 8x22B MOE Model

twitter.com

31–40 of 161 posts

Re: Mistral AI Launches New 8x22B MOE Model

#31
post #28

To this day 8x7b Mixtral remains the best model you can run on a single 48GB GPU. This has the potential to become the best model you can run on two such GPUs, or on an MBP with maxed out RAM, when 4-bit quantized.

My first thought was how much RAM? Will it work on 64GB M1?

Nope. Just the weights would take 88GB at 4 bit. 128GB MBP ought to be able to run it. If I were to guess, a version for Apple MLX should be available within a few days, for those of us fortunate enough to own such a thing.

Re: Mistral AI Launches New 8x22B MOE Model

#33
post #28

To this day 8x7b Mixtral remains the best model you can run on a single 48GB GPU. This has the potential to become the best model you can run on two such GPUs, or on an MBP with maxed out RAM, when 4-bit quantized.

My first thought was how much RAM? Will it work on 64GB M1?

It is ~260GB with presumably fp16 weights. Should fit into 64GB at 3-bit quantization (~49GB).

Edit: To add to this, I've had good luck getting solid output out of mixtral 8x7b at 3-bit, so that isn't small enough to completely kill the model's quality.

Re: Mistral AI Launches New 8x22B MOE Model

#34

Earlier quoted context omitted.

Not sure trying to download the torrent and checking it out

For those of us without twitter, how many GB is the model?

(hope this isn't against rules but) If you don't have Twitter, the magnet link is

  magnet:?xt=urn:btih:9238b09245d0d8cd915be09927769d5f7584c1c9&dn=mixtral-8x22b&tr=udp%3A%2F%http://2Fopen.demonii.com%3A1337%2Fannounce&tr=http%3A%2F%http://2Ftracker.opentrackr.org%3A1337%2Fannounce

Re: Mistral AI Launches New 8x22B MOE Model

#39
post #7

Earlier quoted context omitted.

I've heard command-r is first opensource to beat gpt4 in benchmarks

It beats the old GPT4 version in lmsys benchmark you can check it out here https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... but Command R is commercially licensed We can assume that mistral will do a better job.

> but Command R is commercially licensed

It is licensed under CC-BY-NC-4.0. That license means you are free to use, modify and redistribute it, so long as you aren't doing so "commercially". What exactly counts as "commercial" use is a complex legal question, and the answer may vary from jurisdiction to jurisdiction (different courts may interpret the phrase differently). But, for example, if you are just using it at home for private experimentation on your own personal time, with no plans to make money from doing so (whether now or in the future), I think pretty much everyone will agree that counts as "non-commercial".

Other cases – e.g., if a government agency uses the software to provide some government function, is that "non-commercial"? – are far less clear. Those are really the kind of questions you need to ask a lawyer (which I am not).

Re: Mistral AI Launches New 8x22B MOE Model

#40

Why are some of their models open, and others closed? What is their strategy?

Mistral have stated they want to chase the fine-tune dollar to support le research. We should get thrown a bone of hard to tune mid-range stuff occasionally. Especially when big announcements about small models are expected later in the week (llama3) or when haiku is stealing the thunder from mixtral 8x7b.
Post reply on HN