Live data from Hacker News

Why Ratings Systems Don't Work

goodfil.ms

41–50 of 165 posts

Re: Why Ratings Systems Don't Work

#41
Scatter plots are definitely more informative, once one gives them a couple of minutes to get used to them. However, I think you're shooting at the wrong target and your solution would exacerbate the root problem: bias.

The first time I really noticed the problem was when I published my own Flash game on Kongregate and started paying closer attention to the ratings. That led me to examine my own rating habits and I conjectured that is probably what happens to everyone else.

The bias I'm talking about is caused by the fact that most people can't be bothered to rate something. Most people only rate something when there's a powerful impulse to do so, so most of the votes will be 5 stars or 1 star. The 4-star ratings come from people who liked something enough to be moved to rate it, but not enough to gush about it; note that the group of people who makes that distinction is already substantially smaller than the 5- and 1-star reviewers. The rest comes from a very small minority, most of whom are people who didn't have anything better to do at that moment and decided to spend some time rating, but don't do it on the regular basis.

By the way, I realize that this is just a conjecture, but from what I've seen so far, it seems to be pretty accurate.

I think that introducing an additional axis will only exacerbate this, by raising the bar for rating. If the act of rating starts demanding more effort, you'll get a distribution that is even more skewed than now.

The two improvements I would like to see are:

1. a system that infers ratings from users' actions

2. better mechanisms for gauging the relevance of someone's review/rating based on my preferences/tastes

The first would help reduce the bias and the second would help me extract more useful information from the biased dataset.

Re: Why Ratings Systems Don't Work

#42
post #24

Why would someone who is looking for a film care about rewatchability? Presumably they'll be watching it for the first time and then can decide for themselves whether to ever watch it again. Or did I miss the point and rewatchability is just a placeholder for something more useful?

Article says about Starship Trooper :

there’s a lot of disagreement over whether it’s high quality of not, but generally this scores high-rewatchability. So, maybe not the most intelligent movie, but good fun.

What I intuitively deduced from this example is that rewatchability is metric of enjoyability.

On a single 5 stars rating, some people will give 5 stars because they really really enjoyed the movie, some others will give 5 stars because they thought the film was perfect on a cinematrographic-quality (i.e. scenario, cinematography, acting, casting, etc. insert here some academy-award-technical-category) point of view.

Re: Why Ratings Systems Don't Work

#43
post #12

Histograms are useful int the case of items with a strong love/hate split. The canonical example is the SICP ratings on Amazon: 3.5 average; 177 ratings, 96 five stars, 53 one stars. http://www.amazon.com/Structure-Interpretation-Computer-Prog...

Had to go look up that book. It's actually available free from MIT: http://mitpress.mit.edu/sicp/full-text/book/book.html

Re: Why Ratings Systems Don't Work

#45
I definitely go by the 4.5 stars == very good, E.g. when I go to Amazon I don't buy some random product with a 4.5 star review -- I search for a specific product or a specific kind of product and then reject candidates which are lousy. How is that not INCREDIBLY useful? Similarly, who goes to a movie simply based on whether it's good or not.

In general, if you create any point rating system people who like a thing will tend to rate it towards the top of the scale, e.g. 4/5 or 9/10.

I actually did an informal experiment -- I used to run role-playing tournaments, and do exit surveys on participants. For the first few years we asked players to rate us on a 5-point scale and scored slightly over 4/5 on average. Then we switched to a 10-point scale and scored slightly over 9/10. Not scientific -- but I don't think we suddenly got better.

This finding is backed up by serious research (which is why when a psychologist creates a scale, the numerical ranges need to stay constant in follow-up studies or the results are not statistically comparable).

Netflix, which tries to give users customized ratings, actually subtracts value (in my opinion) from its scores because it tries to make ratings mean "how much will you enjoy this?" BZZZT. I pick stuff for me, my wife, my au pair, and my kids. We don't all like the same stuff, and we don't want to track ratings individually. My kids want good kid stuff. I want good me stuff. Don't try to guess what I like based on our collective tastes.

Re: Why Ratings Systems Don't Work

#47
The old Latin proverb "Quis custodiet ipsos custodes?"

http://en.wikipedia.org/wiki/Quis_custodiet_ipsos_custodes%3...

might in this context be paraphrased to "Who is rating the raters?" The hope in any online rating system is that enough people will come forward to rate something that you care about so that the people who have crazy opinions will be mere outliers among the majority of raters who share your well informed opinions. But how do you ever know that when you see an online rating of something that you haven't personally experienced?

Amazon has had star ratings for a long time. I largely ignore them. I read the reviews. For mathematics books (the thing I shop for the most on Amazon), I look for people writing reviews who have read other good mathematics books and who compare the book I don't know to books I do know. If an undergraduate student whines, "This book is really hard, and does a poor job of explaining the subject" while a mathematics professor says, "This book is more rigorous than most other treatments of the subject," I am likely to conclude that the book is a good book, ESPECIALLY if I can find comments about it being a good treatment of the subject on websites that review several titles at once, as for example websites that advise self-learners on how to study mathematics.

The problem with any commercial website with ratings (Amazon, Yelp, etc., etc.) is that there is HUGE incentive to game the ratings. Authors post bad ratings for books by other authors. The mother and sister and cousins of a restaurant owner post great ratings for their relative's restaurant, and lousy ratings for competing restaurants. I usually have no idea what bias enters into an online rating. So I try to look for the written descriptions of the good or service being sold, and I try to look for signals that the rater isn't just making things up and really knows what the competing offerings are like. When I am shopping for something, I ask my friends (via Facebook, often enough) for their personal recommendations of whatever I am shopping for. Online ratings are hopelessly broken, because of lack of authentication of the basis of knowledge of the raters, so minor details of dimensions of rating or of data display are of little consequence for improving online ratings.

Re: Why Ratings Systems Don't Work

#48
post #34

Earlier quoted context omitted.

If I have already seen this movie why would I need ratings of any kind?

What use case are you referring to? I don't know anyone who uses ratings sites after seeing a movie.

My point exactly. And why would you care about rewatchability if you didn't see the movie once?

Re: Why Ratings Systems Don't Work

#49
post #8

Terrible article. Calling histograms awful, based on nothing more than an opinion. Then trying to conclude that some convoluted scatter plot system makes more sense is laughable. Not to mention, this system is still just a star rating system. This would be no different than having two histograms side by side.. assuming, of course, that you'd even want to rate different aspects of the same thing. I can't even imagine…

This would be no different than having two histograms side by side.. Yes it would, and the article shows why and how. Scatter plots are easy to read (for comp./math. educated people). Two histograms side by side are easy to read (and find correlations) for nobody.

I think it's a problem of brevity. You write something, you refine, you cut out the pet witty comment. You cut your message to the bone. Now you're done.

Read the text of the article describing each movie. The correlations are absolutely meaningless. He doesn't hit on them at all. In fact, with comments like "almost nobody" he's specifically looking at averages. Then coming to a conclusion effectively based on the average of Score 1 and the average of Score 2 to determine what type of movie it is.

So on this Quality/Rewatchability grading system:

Starship Troopers: 3:4 The Fifth Element: 4:4 Blade Runner: 4.5:3.8

Drop that in place of the graphs and the conclusions would sound as seemingly valid.

Re: Why Ratings Systems Don't Work

#50

Terrible article. Calling histograms awful, based on nothing more than an opinion. Then trying to conclude that some convoluted scatter plot system makes more sense is laughable. Not to mention, this system is still just a star rating system. This would be no different than having two histograms side by side.. assuming, of course, that you'd even want to rate different aspects of the same thing. I can't even imagine…

Terrible comment. I found the article interesting. As someone who has played around with many different types of rating systems, I applaud their effort at trying something different. Sounds like you area little too emotionally invested in histograms. I'm not even going to ask why.
Post reply on HN