Live data from Hacker News

Mistral AI Launches New 8x22B MOE Model

twitter.com

151–160 of 161 posts

Re: Mistral AI Launches New 8x22B MOE Model

#151

Earlier quoted context omitted.

Well the motherboard and CPU can be had for $1450. As they're built around standard cases and power supplies and storage, many folks like me will have those already - far less costly than buying the same from Apple if you don't. Spend what you want on ram, unlike with Apple, you can upgrade it any time. Can't reuse my old parts on a brand new Mac, or upgrade it later if I find I need more. Lock-in is rough. https://w…

Note that this is a "QS" CPU, very likely B0 stepping ES by posts elsewhere. A new one of those is around 3k USD alone. The 16-core can be had for around $1200 however and the board for $780. 12x32=384 GB of RAM seems to be about $1400 right now. Going for less capacity don't save that much, unlike the insanely marked up apple memory. And then you need the CPU heatsink for $130.

The errata isn't particularly scary: https://www.amd.com/content/dam/amd/en/documents/epyc-techni...

Re: Mistral AI Launches New 8x22B MOE Model

#152

Earlier quoted context omitted.

I’ve been testing a lot of LLMs on my MacBook and I would say that all of them are far away from being as good as GPT-4, at any time. Many are as good as GPT-3 though. There are also a lot of models that are fine tuned for specific tasks. Language support is one big thing that is missing from open models. I’ve only found one model that can do anything useful with Norwegian, which has never been an issue GPT-4.

Which ones have you tested? There were some huge ones released recently.

Samantha, llama 2 pubmed, marcoroni, openchat, fashiongpt, falcon 180B, deepseek llm chat, orca 2, orca 2 alpac uncersored, meditron, tigerbot, mixtral instruct, wizardcoder, gemma, nouse hermes 2 solar, yarn solar 64k, nouse hermes 2 yi, nous hermes 2 mixtral, nouse hermes llama 2, starcode2, hermes 2 pro mistral, norskgpt mistral and norskgpt llama.

Nouse Hermes 2 Solar is the best model for Norwegian that I've tried so far. It's much better than NorskGPT Mistral/Llama. I actually got it to make fairly decent summaries of news articles, though it wouldn't follow any stricter commands like producing 5 keywords in a json list. Kept producing more than 5 keywords and if I doubled down on the restriction on the number of keywords it would start messing up the json.

The best competitor to GPT-4 was falcon 180b, it's still terrible compared to GPT-4. Mixtral is my new favourite though, it's faster than falcon and in general as good or better. Though I would still pick GPT-4 over Mixtral any day of the week, it's leagues ahead of Mixtral.

Tigerbot has a very interesting trait. It tends to disagree when you try to convince it that it's wrong.

I haven't been able to test out the new 8x22 mixtral or command r plus. These are the next ones on my list!

Re: Mistral AI Launches New 8x22B MOE Model

#153

Earlier quoted context omitted.

Do you have any feel for the performance compared to the M3 Max?

LLM inference is mostly memory bound. An 12-channel Epyc Genoa with 4800MT/s DDR5 ram clocks at 460.8 GB/sec. It's more than the 400GB/s of the M3 Max, only part of that accessible for the CPU. And on the Epyc System you can plug much more memory for when you need larger memory and PCI-E gpus, for when you need less faster memory. Threadripper PRO is only 8-channel, but with memory overclocking it might reach numbers…

That's interesting. It's about the same speed as the M3 Max then.

Have you tested it yourself?

Re: Mistral AI Launches New 8x22B MOE Model

#154

Earlier quoted context omitted.

Which ones have you tested? There were some huge ones released recently.

Samantha, llama 2 pubmed, marcoroni, openchat, fashiongpt, falcon 180B, deepseek llm chat, orca 2, orca 2 alpac uncersored, meditron, tigerbot, mixtral instruct, wizardcoder, gemma, nouse hermes 2 solar, yarn solar 64k, nouse hermes 2 yi, nous hermes 2 mixtral, nouse hermes llama 2, starcode2, hermes 2 pro mistral, norskgpt mistral and norskgpt llama. Nouse Hermes 2 Solar is the best model for Norwegian that I've tri…

Just tested out Command R+ with some niche SHACL constraint questions and it performs considerably worse than GTP-4. Might be a bit better than GPT-3.5 though, which is actually pretty amazing.

Re: Mistral AI Launches New 8x22B MOE Model

#155

Earlier quoted context omitted.

LLM inference is mostly memory bound. An 12-channel Epyc Genoa with 4800MT/s DDR5 ram clocks at 460.8 GB/sec. It's more than the 400GB/s of the M3 Max, only part of that accessible for the CPU. And on the Epyc System you can plug much more memory for when you need larger memory and PCI-E gpus, for when you need less faster memory. Threadripper PRO is only 8-channel, but with memory overclocking it might reach numbers…

That's interesting. It's about the same speed as the M3 Max then. Have you tested it yourself?

Should be at least twice the speed of the M3 Max, as the M3 CPU or GPU only get about half the memory bandwidth available to the package each. M3 Max can't take full advantage of it's memory bandwidth unless CPU, GPU, and NPU are all working at the same time.

Re: Mistral AI Launches New 8x22B MOE Model

#157

Earlier quoted context omitted.

That's interesting. It's about the same speed as the M3 Max then. Have you tested it yourself?

Should be at least twice the speed of the M3 Max, as the M3 CPU or GPU only get about half the memory bandwidth available to the package each. M3 Max can't take full advantage of it's memory bandwidth unless CPU, GPU, and NPU are all working at the same time.

I tried looking for some info on this but could only find the M1 Max review over at anandtech that managed to push 200 GB/s when using multiple cores on the CPU, but couldn’t really get any numbers for just the GPU that seemed realistic.

Do you have a source for the GPU only having access to half the bandwidth of the memory?

Re: Mistral AI Launches New 8x22B MOE Model

#158

Earlier quoted context omitted.

Note that this is a "QS" CPU, very likely B0 stepping ES by posts elsewhere. A new one of those is around 3k USD alone. The 16-core can be had for around $1200 however and the board for $780. 12x32=384 GB of RAM seems to be about $1400 right now. Going for less capacity don't save that much, unlike the insanely marked up apple memory. And then you need the CPU heatsink for $130.

The errata isn't particularly scary: https://www.amd.com/content/dam/amd/en/documents/epyc-techni...

I'm only seeing the errata for the B1 stepping there, not the B0 stepping that are those "QS".

Re: Mistral AI Launches New 8x22B MOE Model

#159

Earlier quoted context omitted.

The errata isn't particularly scary: https://www.amd.com/content/dam/amd/en/documents/epyc-techni...

I'm only seeing the errata for the B1 stepping there, not the B0 stepping that are those "QS".

Good catch! I'd missed that. Still, experiences of folks in the level1tech forums are positive: https://forum.level1techs.com/t/genoa-9654-qs-experiences/19...

Re: Mistral AI Launches New 8x22B MOE Model

#160

Earlier quoted context omitted.

Samantha, llama 2 pubmed, marcoroni, openchat, fashiongpt, falcon 180B, deepseek llm chat, orca 2, orca 2 alpac uncersored, meditron, tigerbot, mixtral instruct, wizardcoder, gemma, nouse hermes 2 solar, yarn solar 64k, nouse hermes 2 yi, nous hermes 2 mixtral, nouse hermes llama 2, starcode2, hermes 2 pro mistral, norskgpt mistral and norskgpt llama. Nouse Hermes 2 Solar is the best model for Norwegian that I've tri…

Just tested out Command R+ with some niche SHACL constraint questions and it performs considerably worse than GTP-4. Might be a bit better than GPT-3.5 though, which is actually pretty amazing.

You need to use their beginning and end token scheme and set rep pen to 1 to get good quality out of cr+.
Post reply on HN