Earlier quoted context omitted.
ah, well there's actually two classes of replies and maybe i'm confusing one for the other here. My claim regarding architecture follows just formally: you can take any statistical model trained via gd and phrase it as a kNN. The only difference is how hard it is to produce such a model from fitting to data, rather than from rephrasing. The idea that there's something special about architecture is, really, a hardware…
I think I see the crux of the disagreement. > The idea that there's something special about architecture is, really, a hardware illusion. Any empirical function approximation algorithm, designed to find the same conditional probability structure, will in the limit t->inf, approximate the same structure (ie., the actual conditional joint distribution of the data). But it's not just about hardware. Maybe it would be, i…
transformers, certainly, arent "informative" in this sense: they start with no prior model of how text would be distributed given the structure of the world.
these arguments all make radical assumptions that we are in somethihng like a physics experiment -- rather than scraping glyphs from books and replaying their patterns