Earlier quoted context omitted.
> Doesn't Language itself encode multimodal experiences Of course it does. We immediately encode pictures/words/everything into vectors anyway. In practice we don't have great text datasets to describe many things in enough detail, but there isn't any reason we couldn't.
There are absolutely reasons that we cannot capture the entirety—or even a proper image—of human cognition in semantic space. Cognition is not purely semantic. It is dynamic, embodied, socially distributed, culturally extended, and conscious. LLMs are great semantic heuristic machines. But they don't even have access to those other components.
You are conflating the embedding layer in an LLM and an embedding model for semantic search.