Also see A Better Lesson: https://rodneybrooks.com/a-better-lesson/
I think the convergence to transformer architectures already proofs this article from 2019 wrong in some ways. Even though CNNs did and do a good job, they will most likely be superseded by vision transformers. Of course you could argue that stable diffusion models and LLMs also have some domain knowledge inside. But it might be a little less with multimodal networks.
So I guess the lesson should be: Don't put more domain knowledge inside than what's needed to make a product out of it right now, because everything else won't progress the field.