Earlier quoted context omitted.
> and convince a hundred million people to install your software to feed you copious real-world training data that you can use to improve model performance. That's actually the easy part. You already have the music. Distorting it by superimposing background noise is really not difficult.
Lol. When you superimpose noise, the original data is still there. When you have a FM radio playing staticky, heavily compressed music through crappy speakers in an acoustically terrible store and being captured by a terrible microphone and then being compressed, a significant amount of nonlinear distortion has taken place. That is extremely hard to model. And you would have to model it or have real data to train a n…
I think you just confirmed how easy (and cheap) it is to actually generate this data.