I appreciate the video and generally agree with Fei-Fei but I think it almost understates how different the problem of reasoning about the physical world actually is. Most dynamics of the physical world are sparse, non-linear systems at every level of resolution. Most ways of constructing accurate models mathematically don’t actually work. LLMs, for better or worse, are pretty classic (in an algorithmic information t…
Regarding sparse, nonlinear systems and our ability to learn them: There is hope. Experimental observation is, that in most cases the coupled high dimensional dynamics almost collapses to low dimensional attractors. The interesting thing about these is: If we apply a measurement function to their state and afterwards reconstruct a representation of their dynamics from the measurement by embedding, we get a faithful r…
Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
131–140 of 163 posts
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#132I've always wondered how spatial reasoning appears to be operating quite differently from other cognitive abilities, with significant individual variations. Some people effortlessly parallel park while others struggle with these tasks despite excelling at other forms of pattern recognition. What was particularly intriguing for me is that some people with aphantasia have no difficulty with spatial reasoning tasks, so…
my theory is that aphantasia is purely about conscious access to visualizing not the existence of the ability to visualise. I have aphantasia but I would say that spatial reasoning is one of the things my brain is the best at
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#133Earlier quoted context omitted.
> We know that universal solutions can’t exist and that all practical solutions require exotic high-dimensionality computational constructs that human brains will struggle to reason about. This has been the status quo since the 1980s. This particular set of problems is hard for a reason. This made me a bit curious. Would you have any pointers to books/articles/search terms if one wanted to have a bit deeper look on t…
I'm not aware of any convenient literature but it is relatively obvious once someone explains it to you (as it was explained to me). At its root it is a cutting problem, like graph cutting but much more general because it includes things like non-trivial geometric types and relationships. Solving the cutting problem is necessary to efficiently shard/parallelize operations over the data models. For classic scalar data…
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#134I've always wondered how spatial reasoning appears to be operating quite differently from other cognitive abilities, with significant individual variations. Some people effortlessly parallel park while others struggle with these tasks despite excelling at other forms of pattern recognition. What was particularly intriguing for me is that some people with aphantasia have no difficulty with spatial reasoning tasks, so…
I have had this idea about parking a car... Most people have proprioception - you know where the parts of your body are without looking. Close your eyes and you intuitively know where your hands and fingers are. When parking a car, it helps to sort of sit in the drivers seat and look around the car. Turn your neck and look past the back seat where your rear tire would be. sense the edges of the car. I think if you so…
It's made me realize that objects are much further from the boundaries of my car when backing into a spot parallel parking. I would never think to get so close to another car if I had to only rely on my own senses.
With that said, I realize there's a significant number of people that are even poorer estimators of these distances than myself. I.e. those that won't pass through two cars even though to me it's obvious that they could easily pass.
I have to imagine a big part of this has to do with risk assessment and lack of risk-free practice opportunity IRL. Nobody is seeing how far they can push or train themselves in this regard when the consequences are to scratch up your car and others' cars. With the birdseye view I can actually do that now!
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#135I appreciate the video and generally agree with Fei-Fei but I think it almost understates how different the problem of reasoning about the physical world actually is. Most dynamics of the physical world are sparse, non-linear systems at every level of resolution. Most ways of constructing accurate models mathematically don’t actually work. LLMs, for better or worse, are pretty classic (in an algorithmic information t…
Bill Peebles is right, naturalistic, physical laws can be learned in deep neural nets from videos.
OR
Fei-Fei Li is right, you need 3D point cloud videos.
Okay, if you think Bill Peebles is right, then all this stuff you are talking about doesn't matter anymore. Lots of great reasons Bill Peebles is probably right, biggest reason of all is that Veo, Sora etc. have really good physics understanding.
If you think Fei-Fei Li is right, you are going to be augmenting real world data sets with game engine content. You can exactly create whatever data you need, for whatever constraints, to train performantly. I don't think this data scalability concern is real.
A compelling reason you are wrong and Fei-Fei Li's specific bet on scalability is right is the existence of Waymo and Zoox. There are also NEW autonomous vehicle companies achieving things faster than Zoox and Waymo did, because a lot of spatial intelligence problems are actually regulatory/political, not scientific.
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#136Earlier quoted context omitted.
my theory is that aphantasia is purely about conscious access to visualizing not the existence of the ability to visualise. I have aphantasia but I would say that spatial reasoning is one of the things my brain is the best at
How does one determine they have aphantasia? How do you know that you are not doing exactly this thing people call visualizing when you perform spatial reasoning?
https://twistedsifter.com/wp-content/uploads/2023/10/AppleVi...
I can only assume people are trying to accurately describe their own experience so when my experience seems to differ a lot it seems to me that there is more going on than just confusion about wording.
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#137I'm particularly hung up on the data problem she touched on (41 min). She rightly points out that unlike language, where we could bootstrap LLMs with the vast, pre-existing corpus of the internet, there's no equivalent "internet of 3D space." She mentions a "hybrid approach" for World Labs, and that's where the real engineering challenge seems to lie.
My mind immediately goes to the trade-offs. If you lean heavily on synthetic data, you're in a constant battle with the "sim-to-real" gap. It works for narrow domains, but for a general "world model," the physics, lighting, and material properties have to be perfect, which is a monumental task. If you lean on real-world capture (e.g., massive-scale photogrammetry, NeRFs, etc.), the MLOps and data pipeline challenges seem staggering. We're not just talking text files; we're talking about petabytes of structured, multi-sensor data that needs to be processed, aligned, and labeled. It feels like an entirely new class of data infrastructure problem.
Her hiring philosophy of "intellectual fearlessness" (31 min) makes a lot of sense in this context. You'd need a team that's not intimidated by the fact that the foundational dataset for their entire field doesn't even exist yet. They have to build the oil refinery while also figuring out where to drill for oil.
It's exciting to see a team with this much deep learning and computer vision firepower aimed at such a foundational problem. It pulls the conversation away from just optimizing existing architectures and towards creating entirely new categories. It leaves me wondering: what does the "AlexNet moment" for spatial intelligence even look like? Is it a novel model architecture, or is the true breakthrough a new form of data representation that makes this problem tractable at scale?
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#138Earlier quoted context omitted.
yeah, see my other comment. To me its totally obvious that we will have a plethora of very valuable startups who use RL techniques to solve realworld problems in practical areas of engineering .. and I just get blank stares when I talk about this :] Ive stopped saying AI when I mean ML or RL .. because people equate LLMs with AI. We need better ML / RL algos for CV tasks : - detecting lines from pixels - detecting ge…
You said : "- detecting lines from pixels - detecting geometry in pointclouds - constructing 3D from stereo images, photogrammetry, 360 panoramas" ==> For me it is more something like : Source = crude video-or-photo pixels (to) ===> Find simple many rectangle-surface that are glued together one another. This is, for me, how you really go easily to detecting rather complexes geometry of any room.
Similarly I use another algo to detect pipe runs which tend to appear as half cylinders in the pointcloud, as the scanner usually sees one side, and often the other side is hidden, hard to access, up against a wall.
So, I guess my point is the devil is in the details .. and machine learning can optimize even further on good heuristics we might come up with.
Also, when you go thru a whole pointcloud, you have a lot of data to sift thru, so you want something fairly efficient, even if your using multiple GPUs do do the heavy matmull lifting.
You can think of RL as an optimization - greatly speeding up something like monte carlo tree search, by learning to guess the best solution earlier.
Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#139makes sense - humans have evolved a lot of wetware dedicated to 3D processing from stereo 2D. I've made some progress on a PoC in 3D reconstruction - detecting planes, edges, pipes from pointclouds from lidar scans, eg : https://youtu.be/-o58qe8egS4 .. and am bootstrapping with in-house gigs as I build out the product. Essentially it breaks down to a ton of matmulls, and I use a lot of tricks from pre-LLM ML .. this…
Have you tried "traditional" approaches like a Delaunay triangulation on the point cloud, and how does your method compare to that? Or did you encounter difficulties with that? Regarding what you say of planes and compression, you can look into metric-based surface remeshing. Essentially, you estimate surface curvature (second derivatives) and use that to distort length computations, remeshing your surface to length…
Efficient re-meshings are important, and its worth improving on the current algorithms to get crisper breaklines etc, but you really want to go a step further and do what humans do manually now when they make a CAD model from a pointcloud - ie. convert it to its most efficient / compressed / simple useful format, where a wall face is recognized as a simple plane. Even remeshing and flat triangle tesselation can be improved a lot by ML techniques.
As with pointclouds, likewise with 'photogrammetry', where you reconstruct a 3D scene from hundreds of photos, or from 360 panoramas or stereo photos. I think in the next 18 months ML will be able to reconstruct an efficient 3D model from a streetview scene, or 360 panorama tour of a building. An optimized mesh is good for visualization in a web browser, but its even more useful to have a CAD style model where walls are flat quads, edges are sharp and a door is tagged as a door etc.
Perhaps the points Im trying to make are :
- the normal techniques are useful but not quite enough [ heuristics, classical CV algorithms, colmap/SfM ]
- NeRFs and gaussian splats are amazing innovations, but dont quite get us there
- to solve 3D reconstruction, from pointclouds or photos, we need ML to go beyond our normal heuristics : 3D reality is complicated
- ML, particularly RL, will likely solve 3D reconstruction quite soon, for useful things like buildings
- this will unlock a lot of value across many domains - AEC / construction, robotics, VR / AR
- there is low hanging fruit, such as my algo detecting planes and pipes in a pointcloud
- given the progress and the promise, we should be seeing more investment in this area [ 2Mn of investment could potentially unlock 10Bn/yr in value ]Re: Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
#140It's hard to describe, but it's felt like LLMs have completely sucked the entire energy out of computer vision. Like... I know CVPR still happens and there's great research that comes out of it, but almost every single job posting in ML is about LLMs to do this and that to the detriment of computer vision.
On the other hand I just chatted with Opus 4 for the first time a few minutes ago and I am completely blown away.