Live data from Hacker News

Mixtral 8x22B

mistral.ai

11–20 of 252 posts

Re: Mixtral 8x22B

#14
"64K tokens context window" I do wish they had managed to extend it to at least 128K to match the capabilities of GPT-4 Turbo

Maybe this limit will become a joke when looking back? Can you imagine reaching a trillion tokens context window in the future, as Sam speculated on Lex's podcast?

Re: Mixtral 8x22B

#16
post #8

I just find it hilarious how approximately 100% of models beat all other models on benchmarks.

What do you mean?

Virtually every announcement of a new model release has some sort of table or graph matching it up against a bunch of other models on various benchmarks, and they're always selected in such a way that the newly-released model dominates along several axes.

It turns interpreting the results into an exercise in detecting which models and benchmarks were omitted.

Re: Mixtral 8x22B

#17
post #3

Great to see such free to use and self-hostable models, but it's said that open now means only that. One cannot replicate this model without access to the training data.

...And a massive pile of cash/compute hardware.

not that massive, we're talking six figures. There was a blogpost about this a while back on the startpage of HN.

Re: Mixtral 8x22B

#20
I'm really excited about this model. Just need someone to quantize it to ~3 bits so it'll run on a 64GB MacBook Pro. I've gotten a lot of use from the 8x7b model. Paired with llamafile and it's just so good.
Post reply on HN