This is stunning! Great stuff. Since the input and prediction is a single sequence, did you experiment with beamsearch/stochastic beamsearch decoding (maybe with additional diversity criteria)? I found that even simple models (markov chains) got a big diversity boost with a stochastic beamsearch - it might avoid the problems with low temperature repetition that could happen in a standard beamsearch. However, my music…
There's also no consensus on whether the high- or low-temperature samples sound better. I've heard both opinions from several people.
Sageev did the final rendering, not sure what he used but I'm pretty sure it was nothing too fancy.