Does anyone know if these models can output also Midi instead of plain audio?
However, there are many models which do output midi. That's actually much simpler, and has been done already a few years ago.
I thought OpenAI did this. But then, I might misremember, because their Jukebox actually also seems to produce raw audio (https://openai.com/blog/jukebox/).
Edit: Ah, it was even earlier, OpenAI MuseNet, this: https://openai.com/blog/musenet/
However, midi generation is so easy, you even find it in some tutorials: https://www.tensorflow.org/tutorials/audio/music_generation