Wonder how far off the whole "generate music based on your existing music library" thing is going to be? That'll make musicians happy with big tech as well, just like artists are. *sigh*
Perhaps LoRA (Low-Rank Adaptation) training techniques could be used for these types of models, like they're currently being used with LLMs and latent text-to-image diffusion models.
> Mitigations: Vocals have been removed from the data source using corresponding tags, and then using a state-of-the-art music source separation method, namely using the open source Hybrid Transformer for Music Source Separation (HT-Demucs).
> Limitations: The model is not able to generate realistic vocals.
(https://github.com/facebookresearch/audiocraft/blob/main/mod...)
I suspect this was a combination of playing it safe and that the model isn't well architected to reproduce meaningful vocals.