Maybe I'm totally misreading this, but it seems like the post contradicts itself. At the beginning of the third paragraph: > Impressively, open source models have been able to quickly catch up to big labs. And then the beginning of the fourth: > Open-source has been lagging behind proprietary models for years, but lately this gap has been widening. Followed by a picture that is more or less inscrutable.
Hey, I'm the author of the post. The image has been fixed, and the point I'm making is that proprietary models are almost always ahead, and this gap is widening. OS models that are nearly at the same quality are usually distilled versions of proprietary models, or somehow get training data from them. Sometimes, after massive, expensive training runs models are open sourced anyway, and at some point that becomes unsus…
I think I'm getting it now: OS models are getting closer, but only via distillation. Not by training a new frontier model which is out of reach for economic reasons.