Earlier quoted context omitted.
Yes, you could definitely gather non-anecdotal, statistical evidence, but I think you're missing dschobel's main point. I think he's making a point about the inherently subjective nature of "quality of recommendation", not suggesting the obviously stupid idea that it's impossible to gather statistical data on a recommendation engine. And to that point, I have to agree with him. At least, most of the obvious choices f…
How did Netflix evaluate submissions to its competition?
I thought about the Netflix example, which is obviously related, while writing my post. Ultimately the Netflix competition isn't for a recommendation engine, it's a prediction engine. They use that prediction engine to produce recommendations, but that's a separate phase. The top N predicted ratings != the best N things to recommend at this moment, though obviously it's useful to know predicted ratings for movies if you're writing a recommender.
A recommendation engine should not just recommend the highest predicted rated apps given that the user downloads the app. As an extremely simple example, even if you habitually rate games much higher than other apps, it shouldn't recommend you 100% games. You probably want to see other things sometimes too. The perfect set of apps on your phone would not be entirely games; you do like to twitter after all, even if you're not entirely fond of any of the twitter apps out right now.
The perfect recommendation engine would be psychic; it would know the set of apps on your phone that would make you maximally happy. That's obviously not the same as predicting exactly what you'd rate each app (though presumably you could do that easily).