I just find it hilarious how approximately 100% of models beat all other models on benchmarks.
Mixtral 8x22B
21–30 of 252 posts
Re: Mixtral 8x22B
#22The development never stops. In a few years we will look back and see how the previous models were and how they're now. How we couldn't run LLaMa 70B on MacBook Air and now we can.
Around 5 years ago, it took a lambda user some pretty significant hardware, software and time (around a full night), to try to create a short deepfake. Now, you don't need any fancy hardware and you can have some decent results within 5 min on your average computer.
Re: Mixtral 8x22B
#23"64K tokens context window" I do wish they had managed to extend it to at least 128K to match the capabilities of GPT-4 Turbo Maybe this limit will become a joke when looking back? Can you imagine reaching a trillion tokens context window in the future, as Sam speculated on Lex's podcast?
Re: Mixtral 8x22B
#24Re: Mixtral 8x22B
#25Re: Mixtral 8x22B
#26Isn't equating active parameters with cost a little unfair since you still need full memory for all the inactive parameters?
Re: Mixtral 8x22B
#27Re: Mixtral 8x22B
#28Wish I had invested in the extra 32GB for my mac laptop.
Re: Mixtral 8x22B
#29Earlier quoted context omitted.
What do you mean?
Virtually every announcement of a new model release has some sort of table or graph matching it up against a bunch of other models on various benchmarks, and they're always selected in such a way that the newly-released model dominates along several axes. It turns interpreting the results into an exercise in detecting which models and benchmarks were omitted.
The different teams are learning from each other and pushing boundaries; there's virtually no reason for any of the teams to release a model or product that is somehow inferior to a prior one (unless it had some secondary attribute such as requiring lower end hardware).
We're simply not seeing the ones that came up short; we don't even see the ones where it fell short of current benchmarks because they're not worth releasing to the public.
Re: Mixtral 8x22B
#30I'm really excited about this model. Just need someone to quantize it to ~3 bits so it'll run on a 64GB MacBook Pro. I've gotten a lot of use from the 8x7b model. Paired with llamafile and it's just so good.