google might have a chance to win the LLM battle, to train multi-modal LLM based on youtube videos, which is a huge advantage over Microsoft. OpenAI's Codex model got the reasoning ability through github. LLM could learn so much more with real-life videos in an embodied way.
That would require better speech recognition. Judging by automated subtitles, it's not there yet. And I don't think YouTube videos tend to have better veracity than the web at large. Maybe YouTube videos can be used to train models on how to verbally present information in a compelling way, but I don't see how it would make an AI better at generating written text.
No, I mean ML can recognize real-world objects like human eyes do. They have a raw, direct understanding of the visual "world", an alternative to parse the world facts only by texts.