The astounding thing about Goliath wasn’t that is was a huge leap in performance, it was that the damn thing functioned at all. To this day, I still don’t understand why this didn’t raise more eyebrows. This wasn't something I really dug into in great detail but I remember my surprise back then at how all those merged models and those "expanded" models like Goliath still generated coherent output. IMO those were more…
Even between two models of identical architecture, they may have landed on quite different internal representations if the training data recipe was substantially different.
But it would be fun to experiment with.