Live data from Hacker News

Mistral AI Launches New 8x22B MOE Model

twitter.com

141–150 of 161 posts

Re: Mistral AI Launches New 8x22B MOE Model

#141
post #52

Out of topic but are we now back at the same performance than ChatGPT 4 at the time people said it worked like magic (meaning before the nerf to make it more politically correct but making his performance crash)?

I’ve been testing a lot of LLMs on my MacBook and I would say that all of them are far away from being as good as GPT-4, at any time. Many are as good as GPT-3 though. There are also a lot of models that are fine tuned for specific tasks. Language support is one big thing that is missing from open models. I’ve only found one model that can do anything useful with Norwegian, which has never been an issue GPT-4.

Which ones have you tested? There were some huge ones released recently.

Re: Mistral AI Launches New 8x22B MOE Model

#142
post #87

Earlier quoted context omitted.

Could you recommend one or a few in particular?

The current best open weights model is probably Cohere Command-R+. The memory requirements on it are quite high, though.

I really want to see some benchmarks with performance weighted by energy use. I think Mistral 7B performance to watt would be the leader by a huge margin. On many tasks I get equal performance on zero shot classification tasks on Mistral than in bigger models.

Re: Mistral AI Launches New 8x22B MOE Model

#143
post #127

Earlier quoted context omitted.

It is there, not for all the benchmarks, but for those where it is included, GPT-4 scores much higher. Not surprising since GPT-4 is still state-of-the-art and much bigger. Where Mistral has been particularly impressive is when you take the size of the model into account.

GPT-4 is instruct tuned model, of course it's going to score higher, apples and oranges.

Yeah and the instruct tunes provided by Mistral on other models are pretty great.

Re: Mistral AI Launches New 8x22B MOE Model

#144

Earlier quoted context omitted.

There is a user called The Bloke on hugging face- they release pre quantized models pretty soon after the full size drop. Just watch their page and pray you can fit the 4 bit in your GPU. I’m sure they are already working on it.

TheBloke stopped uploading in January. There are others that have stepped up though.

Oh really? Who else should I be looking at?

That person is a hero, super bummed!

Re: Mistral AI Launches New 8x22B MOE Model

#145

Earlier quoted context omitted.

There is a user called The Bloke on hugging face- they release pre quantized models pretty soon after the full size drop. Just watch their page and pray you can fit the 4 bit in your GPU. I’m sure they are already working on it.

I think 4b for this is support to be over 70GB, so definitely still heavy hardware.

Fucking hell, my A6000 is shy of that and I can’t reasonably justify picking up a second.

Re: Mistral AI Launches New 8x22B MOE Model

#146

What is the excitement around models that arent as good as llama? This is clearly an inferior model that they are willing to share for marketing purposes. If it was an improvement over llama, sure, but it seems like just an ad for bad AI.

Mixtral 7x8b was way better than llama2 70b and used less RAM and compute at the same time. This model is way better than llama.

In fact I would go as far as saying llama2 isn’t that good compared to some of the most recent models.

Re: Mistral AI Launches New 8x22B MOE Model

#147

Earlier quoted context omitted.

TheBloke stopped uploading in January. There are others that have stepped up though.

Oh really? Who else should I be looking at? That person is a hero, super bummed!

TheBloke's grant ran out.

Re: Mistral AI Launches New 8x22B MOE Model

#148

Earlier quoted context omitted.

I would be very curious to see pricing on Epyc systems with terabytes of RAM that cost less than $6k including the RAM...

Well the motherboard and CPU can be had for $1450. As they're built around standard cases and power supplies and storage, many folks like me will have those already - far less costly than buying the same from Apple if you don't. Spend what you want on ram, unlike with Apple, you can upgrade it any time. Can't reuse my old parts on a brand new Mac, or upgrade it later if I find I need more. Lock-in is rough. https://w…

Note that this is a "QS" CPU, very likely B0 stepping ES by posts elsewhere. A new one of those is around 3k USD alone. The 16-core can be had for around $1200 however and the board for $780.

12x32=384 GB of RAM seems to be about $1400 right now. Going for less capacity don't save that much, unlike the insanely marked up apple memory. And then you need the CPU heatsink for $130.

Re: Mistral AI Launches New 8x22B MOE Model

#149

Earlier quoted context omitted.

I'm curious how the newer consumer Ryzens might fare. With LPDDR5X they have >100 GB/s memory bandwidth and the GPUs have been improved quite a bit (16 TFLOPS FP16 nominal in the 780M). There are likely all kinds of software problems but setting that aside the perf/$ and perf/watt might be decent.

Consumer Ryzens only have two-channel memory controllers. Two dual-rank (double sided) DIMMs per channel, which you would need to use to get enough RAM for LLMs, drops the memory bandwidth dramatically -- almost all the way back down to DDR4 speeds.

For consumer Ryzen to pencil out it would require a cluster of APU-equipped machines with the model striped across them. Given say 16GB of model per machine and 60GBps actual memory bandwidth @ $500 it's favorable vs A100s if the software is workable (which my guess is it's not today due to AMD's spotty support). This is for inference, training probably would be too slow due to interconnect overhead.

Re: Mistral AI Launches New 8x22B MOE Model

#150

Earlier quoted context omitted.

I think you are an optimist here. I can barely run mixtral-8x-7B on my M2 Pro 32G Mac, but I am grateful to be able to run it at all.

Which quantization level are you using?

Q2, so not so great. I usually run other models. I would be embarrassed to tell you how long my “ollama list” is.
Post reply on HN