Live data from Hacker News

Spatial intelligence is AI’s next frontier

drfeifei.substack.com

101–110 of 136 posts

Re: Spatial intelligence is AI’s next frontier

#101
post #80

From reading that, I'm not quite sure if they have anything figured out. I actually agree, but her notes are mostly fluff with no real info in there and I do wonder if they have anything figured out besides "collect spatial data" like imagenet. There are actually a lot of people trying to figure out spatial intelligence, but those groups are usually in neuroscience or computational neuroscience. Here is a summary pap…

> From reading that, I'm not quite sure if they have anything figured out. I actually agree, but her notes are mostly fluff with no real info in there and I do wonder if they have anything figured out besides "collect spatial data" like imagenet. Right. I was thinking about this back in the 1990s. That resulted in a years-long detour through collision detection, physically based animation, solving stiff systems of no…

> I used to make the comment, pre-LLM, that we needed to get to mouse/squirrel level intelligence rather than trying to get to human level abstract AI. But we got abstract AI first. That surprised me.

"AI" is not based on physical real world data and models like our brain. Instead, we chose to analyze human formal (written) communication. ("formal": actual face to face communication has tons of dimensions adding to the text representation of what is said, from tone, speed to whole body and facial expressions)

Bio-brains have a model based on physical sensor data first and go from there, that's completely missing from "AI".

In hindsight, it's not surprising, we skipped that hard part (for now?). Working with symbols is what we've been doing with IT for a long time.

I'm not sure going all out on trying to base something on human intelligence, i.e. human neuro networks, is a winning move. I see it as if we had been trying to create airplanes that flap their wings. For one, human intelligence already exists, and when you lean back and manage to look at how we do on small and large problems from an outside perspective it has plenty of blind spots and disadvantages.

I'm afraid if we were to manage a hundred percent human level intelligence AI we will be disappointed. Sure, it will be able to do a lot, but in the end, nothing we don't already have.

Right now that would also just be the abstract parts, I think the "moving the body" physical parts in relation to abstract commands would be the far more interesting part, but since current AI is not about using physical sensor data at all, never mind combining it with the abstract stuff...

Re: Spatial intelligence is AI’s next frontier

#102

Genie 3 (at a prototype level) achieves the goal she describes: a controllable world model with consistency and realistic physics. Its sibling Veo 3 even demonstrates some [spatial problem-solving ability]( https://video-zero-shot.github.io/ ). Genie and Veo are definitely closer to her vision than anything World Labs has released publicly. However, she does not mention Google's models at all. This omission makes the…

Also Gemini ER, which can kinda just work, spatially, in the real world:

https://deepmind.google/models/gemini-robotics/gemini-roboti...

Re: Spatial intelligence is AI’s next frontier

#103

I think a lot of people are really bad at evaluating world models. Feifei is right here that they are multimodal but really they must codify a physics. I don't mean "physics" but "a physics". I also think it's naïve to think this can be done through data alone. I mean just ask a physicist...[0]. But why people are really bad at evaluating them is because the details dominate. What matters here is consistency. We need…

The YouTube video tells a fascinating story. Who would be our Fermi today who can tell the truth and save five years of work, billions of dollars and careers of Ph.D. students?

We wouldn’t expect LLM to review a paper and tell us the truth like Fermi did. That is super-intelligence.

Thanks for sharing.

Re: Spatial intelligence is AI’s next frontier

#104

I think a lot of people are really bad at evaluating world models. Feifei is right here that they are multimodal but really they must codify a physics. I don't mean "physics" but "a physics". I also think it's naïve to think this can be done through data alone. I mean just ask a physicist...[0]. But why people are really bad at evaluating them is because the details dominate. What matters here is consistency. We need…

https://www.webofstories.com/play/freeman.dyson/94

This link has transcript.

Re: Spatial intelligence is AI’s next frontier

#105

Earlier quoted context omitted.

It's not enough by a long shot. Placement isn't related directly to vicarious trial and error, path integrations, sequence generation. There's a whole giant gap between grid cells and intelligence.

>There's a whole giant gap between grid cells and intelligence. Please check this recent article on the state machine in the hippocampus based on learning [1]. The findings support the long-standing proposal that sparse orthogonal representations are a powerful mechanism for memory and intelligence. [1] Learning produces an orthogonalized state machine in the hippocampus: https://www.nature.com/articles/s41586-024-08…

Of course, but the mechanisms “remain obscure”. The entorhinal cortex is but a facet of this puzzle and placement vs head direction etc must be understood beyond mere prediction. There are too many essential parts that are not understood particularly senses and emotion which play the tinkering precursors to evolutionary function that are excluded now as well as the likelyhood that prediction error and prediction are but mistaken precursor computational bottlenecks to unpredictability. Pushing AI into the 4% of a process materially identified as entorhinal is way premature.

This approach simply follows suit with the blundering reverse engineering of the brain in cog sci where material properties are seen in isolation and processes are deduced piecemeal. The brain can only be understood as a whole first. See rhythms of the brain or unlocking the brain.

There’s a terrifying lack of curiosity in the paper you posted, a kind of smug synthetic rush to import code into a part of the brain that’s a directory among directories that has redundancies as a warning: we get along without this.

Your and their view (OSM) is too narrow. eg categorization is baked into the whole brain. How? This is one of 1000s of processes that generalize materially across the entire brain. Isolating "learning" to the allocortex is incredibly misleading.

https://www.cell.com/current-biology/fulltext/S0960-9822(25)...

Re: Spatial intelligence is AI’s next frontier

#106
>>Spatial Intelligence is the scaffolding upon which our cognition is built.

Human cognition isn’t built on abstract reasoning alone. It’s embodied, grounded in sensation.

Evolution didn’t achieve generalization across domains by making brains more symbolic. It did so by making them more integrated by fusing chemical gradients, touch, proprioception, light, sound, temperature, and pressure into one continuous internal narrative.

Intelligence does not seem to be an algorithmic property; it’s a felt coherence across senses. Our reasoning emerges from a complex interaction of sensory information, memory, emotions, and cognitive processing. Sensory completeness is the way forward.

Re: Spatial intelligence is AI’s next frontier

#107

From reading that, I'm not quite sure if they have anything figured out. I actually agree, but her notes are mostly fluff with no real info in there and I do wonder if they have anything figured out besides "collect spatial data" like imagenet. There are actually a lot of people trying to figure out spatial intelligence, but those groups are usually in neuroscience or computational neuroscience. Here is a summary pap…

The question, as always, is: can we get any useful insights from all of that?

Trying to copy biological systems 1:1 rarely works, and copying biological systems doesn't seem to be required either. CNNs are somewhat brain-inspired, but only somewhat, and LLMs have very little architectural similarity to human brain - other than being an artificial neural network.

This functional similarity of LLMs to the human brain doesn't come from reverse engineered details of how the human brain works - it comes from the training process.

Re: Spatial intelligence is AI’s next frontier

#109

From reading that, I'm not quite sure if they have anything figured out. I actually agree, but her notes are mostly fluff with no real info in there and I do wonder if they have anything figured out besides "collect spatial data" like imagenet. There are actually a lot of people trying to figure out spatial intelligence, but those groups are usually in neuroscience or computational neuroscience. Here is a summary pap…

What I personally find amusing is this part:

>3. Interactive: World models can output the next states based on input actions

>Finally, if actions and/or goals are part of the prompt to a world model, its outputs must include the next state of the world, represented either implicitly or explicitly. When given only an action with or without a goal state as the input, the world model should produce an output consistent with the world’s previous state, the intended goal state if any, and its semantic meanings, physical laws, and dynamical behaviors. As spatially intelligent world models become more powerful and robust in their reasoning and generation capabilities, it is conceivable that in the case of a given goal, the world models themselves would be able to predict not only the next state of the world, but also the next actions based on the new state.

That's literally just an RNN (not a transformer). An RNN takes a previous state and an input and produces a new state. If you add a controller on top, it is called model predictive control. The most extreme form I have seen is temporal difference model predictive control (TD-MPC). [0]

[0] https://arxiv.org/abs/2203.04955

Post reply on HN