Earlier quoted context omitted.
>Well, lets get an LLM to translate one of not understood human languages first. You can do both. Models trained to receive and predict audio tokens alongside text are coming. You could just add such data as part of training. >Does that get us any closer to translating it. NO, because without shared context to start building connections and relationships its highly unlikely that were going to find commonality. If you…
the sounds for man, and the word man can be mapped. The concept of man can be mapped from English to Chinese. Given enough of these clues, and a large enough data set it's not shocking that models get decent at inference between languages. But whales? What we are (probably) proposing is putting two languages side by side with NO context and saying "figure it out". Here is a great example of people and language and pe…
1. Many such concepts cannot be directly mapped between different languages especially with distant language pairs.
2. Language Models and Image models even if both are only trained on their respective modality learn structural representations so similar, you can connect them with a simple linear layer. That's it. Entirely different modalities https://arxiv.org/abs/2209.15162
https://arxiv.org/abs/2304.08485
I think you are severely underestimating the extent to which representations can group in neural networks.
They're mammals, they're social animals, they see. You're assuming a level of alienness we don't have the knowledge or understanding to truly ascertain.