Would this approach destroy critical investments in physics- or modeling-based reasoning? I'm all for the task reasoning and the multi-view recognition, based on relevant points. I'm very uncomfortable with the loose world "understanding". The fault model I see is that e.g., this "visual understanding" will get things mostly right: enough to build and even deliver products. However, these are only probabilistic guara…
Having a model that understands physics helps us certify safety. But how much physics is enough? There's a lot to knowing about gravity. You probably don't need orbital dynamics, but you do need to not jump out a second story window.
(Video generation models are an interesting case study: https://www.alphaxiv.org/overview/2501.09038)