Earlier quoted context omitted.
It's just scale. But scale that comes with more than an order of magnitude more expense than the Llama models. I don't see anyone training such a model and releasing it for free anytime soon
I thought it was revealed to be fundamentally ensemblamatic in a way the others weren’t? Using “experts” I think? Seems like it would meet the bar for “secret sauce” to me
Until this paper (https://arxiv.org/abs/2305.14705) indicated they apparently benefit far more from Instruct tuning than dense models, it was mostly a "good on paper" kind of thing.
In the paper, you can see the underperformance i'm talking about.
Flan-Moe-32b(259b total) scores 25.5% on MMLU pre Instruct tuning and 65.4 after.
Flan 62b scores 55% before Instruct tuning and 59% after.