Earlier quoted context omitted.
But your talking about something they are not today, and quite likely we won’t be calling them LLM’s as the architecture is likely to change quite a lot before we reach a point they are comparable to human capabilities.
CLIP, which powers diffusion models, creates a joint embeddings space for text and images. There's a lot of active work on extending these multimodal embedding spaces to audio and video. Microsoft has a paper just a week or so ago showing that llm's with a joint embeddings trained on images can do pretty amazing things, and (iirc) with better days efficiency than a text only model. These things are already here; it's…
Multiple so called modalities doesn’t necessarily address the shortcomings, if anything it just highlights that there are many steps, and each step has typically created significant changes to the prior architecture!