Live data from Hacker News

How Not to Sort by Average Rating (2009)

evanmiller.org

121–130 of 158 posts

Re: How Not to Sort by Average Rating (2009)

#121
post #109

Earlier quoted context omitted.

Good answer, this has a nice property that it can be applied to any reasonably behaved average-based system to get an honest mechanism. For a plain average it is equivalent to the median.

Not actually equivalent to the median. If all the scores are 1s and 5s, then the median will be a 1 or a 5, but the average will be somewhere in between. The problem is slightly easier to understand if we consider the grading scale from 0 to 100. Then every agent, trying to manipulate the score as much as they can toward their "ideal grade", will submit a grade of 0 or 100. The average will converge to the unique num…

Yes, you're right and I was wrong.

Re: How Not to Sort by Average Rating (2009)

#122
post #46

Earlier quoted context omitted.

I've been saying this for years, thumbs up/down is the only system that makes sense to me. Foursquare uses it and I've found their scores to be way more useful than Yelp's. The biggest problem with star ratings is that it's so arbitrary. What is the difference between 3 and 3.5? What is a 1 vs a 2? 3/5 is 60%, that's almost failing when you think about it on a grading scale, if I scored something as a 3/5 I would nev…

To be honest, the fact that 60% is a failing grade is a failure of the grading system, not a fact to take for granted. We've basically lost the entire dynamic range of 0-60% for no good reason.

I would actually say it's often not strict enough. In what serious field is it acceptable to only know, say, 70% of the material? Do you want to drive on a bridge designed by an engineer who only got 70% on their exams? It depends on how the test is structured, really, but unless it was one of those tests designed to bring smart people to their knees, I'd rather not.

Re: How Not to Sort by Average Rating (2009)

#123

Earlier quoted context omitted.

Counterpoint: I almost solely rely on the stars histogram in Yelp (available only on the website, not the app), completely ignoring whatever Yelp's calculated "average" is. If a place has more 5-star ratings than 4-star ratings, it's generally amazing. If it has more 4-star ratings than 5-star ratings, it's generally fine but not something particularly special. Just thumbs up/down would eliminate what is, to me, the…

I used to think the same thing until I realized the most accurate and consistent ratings I use on a regular basis is rotten tomatoes. And they're based on strict thumbs up/ down. It ensures votes hold equal weight and that "extreme polar" voters don't skew things. It also avoids the opposite problem of "everything is neutral" vote unless horrible/incredible. RT also handles high brow and low brow well. You get less v…

I used to religiously research movies on RT - with a lot of success in my mind. With the user rating, the critic rating, and the "top" critic rating, you can infer a surprising amount about who is going to like any given film, and you learn over time where you fall on the critic/top critic/audience graph.

Recently, however, it seems like more (imo undeserving) movies that are "just ok" - like decent, but nothing special, romantic comedies and big blockbusters - are scoring above 90%. I might be being curmudgeonly about it, but I've nearly stopped checking it because it feels like there's no information there. My theory is that this started happening once Roger Ebert died... without such a leader in the field, no one is willing to say they didn't like a film unless it's obviously very bad.

Re: How Not to Sort by Average Rating (2009)

#124
post #46

Earlier quoted context omitted.

John Gruber has been arguing that the only meaningful way to do ratings is a simple thumbs up/thumbs down. I don't necessarily agree, but I see the appeal. I usually don't want ratings, I want the Wirecutter treatment. Sometimes, I know/care enough to really research the topic, in which case star reviews are relatively unhelpful. The rest of the time, I just want someone trustworthy to say "buy this if you want to pa…

I've been saying this for years, thumbs up/down is the only system that makes sense to me. Foursquare uses it and I've found their scores to be way more useful than Yelp's. The biggest problem with star ratings is that it's so arbitrary. What is the difference between 3 and 3.5? What is a 1 vs a 2? 3/5 is 60%, that's almost failing when you think about it on a grading scale, if I scored something as a 3/5 I would nev…

Perhaps the issue isn't the granularity of a single dimensional rating scale, but the lack of expressive options when in reality your feeling about something is complex and multifaceted.

I've been really interested in the idea of emotive reviews as an alternative to single dimensional scores. The best idea I have at the moment is something akin to emoji reactions like you see on GitHub issues, finding a way to encode some feelings relevant to product reviews in a mechanism like that seems really intriguing to me.

Re: How Not to Sort by Average Rating (2009)

#125
post #46

Earlier quoted context omitted.

I've been saying this for years, thumbs up/down is the only system that makes sense to me. Foursquare uses it and I've found their scores to be way more useful than Yelp's. The biggest problem with star ratings is that it's so arbitrary. What is the difference between 3 and 3.5? What is a 1 vs a 2? 3/5 is 60%, that's almost failing when you think about it on a grading scale, if I scored something as a 3/5 I would nev…

Perhaps the issue isn't the granularity of a single dimensional rating scale, but the lack of expressive options when in reality your feeling about something is complex and multifaceted. I've been really interested in the idea of emotive reviews as an alternative to single dimensional scores. The best idea I have at the moment is something akin to emoji reactions like you see on GitHub issues, finding a way to encode…

I envision a panel of emoticons akin to the Facebook reaction set, but where the user can select as many as they want to quickly convey different combinations of their reactions:

    (thumbs up)      I liked this
    (heart)          I loved this
    (thumbs down)    I didn’t like this
    (smiling face)   This made me happy or satisfied
    (frowning face)  This made me sad or disappointed
    (surprised face) This made me surprised or impressed
    (angry face)     This made me angry or frustrated
Of course, it gets complicated. Did Sam U. Zerr give that product an (angry face) because they used it and didn’t like it, or because they’re offended that you would recommend it, or what?

If you’re only using icons to make recommendations to an individual user based on their own history, maybe you don’t need to infer the actual meanings; you can add all sorts of icons without any particular meaning and just make recommendations by correlation:

    (thinking face)  I’m considering this / I’m confused by or dubious of this
    (gear)           This was useful / this made me think
    (fire)           This album was great / this sauce was spicy
    (heart eyes)     I really want this / this is adorable
    ...
E.g. a recommendation for me might be “(thumbs up)(gear)(heart eyes)” because some product or content is similar, by some hidden metrics, to other things that I’ve reacted to in those ways.

Just brainstorming here. There are obviously many possible approaches in this space.

Re: How Not to Sort by Average Rating (2009)

#126

Earlier quoted context omitted.

To be honest, the fact that 60% is a failing grade is a failure of the grading system, not a fact to take for granted. We've basically lost the entire dynamic range of 0-60% for no good reason.

So you're saying that the measurement of student mastery has no noise floor?

It depends on how things are graded. On a multiple-choice test with four choices per question, someone with no knowledge who guesses randomly will get ~25%. On a true-false test, someone with no knowledge gets ~50%. On a project graded by a human, or a worksheet whose answers are real numbers, someone with no knowledge and a hard-eyed grader might well get 0%. Different classes will have different proportions of these things that contribute to the overall grade (at least, I haven't heard of any requirement that classes have the same proportions of such). The simple approach of summing total points achieved over each graded item, divided by total points possible, is straightforward to calculate, but I think there's no mathematical justification for choosing one percentage-based grading scale and applying it uniformly to all classes.

Re: How Not to Sort by Average Rating (2009)

#127

Earlier quoted context omitted.

Perhaps the issue isn't the granularity of a single dimensional rating scale, but the lack of expressive options when in reality your feeling about something is complex and multifaceted. I've been really interested in the idea of emotive reviews as an alternative to single dimensional scores. The best idea I have at the moment is something akin to emoji reactions like you see on GitHub issues, finding a way to encode…

I envision a panel of emoticons akin to the Facebook reaction set, but where the user can select as many as they want to quickly convey different combinations of their reactions: (thumbs up) I liked this (heart) I loved this (thumbs down) I didn’t like this (smiling face) This made me happy or satisfied (frowning face) This made me sad or disappointed (surprised face) This made me surprised or impressed (angry face)…

Put differently, a set of binary choices: amusing, interesting, sad, ... It's a bit difficult to come up with a good set to rate any thing, but I can see it working for specific topics, like movies or games.

Or, one could just let users tag the subject and the interface would display the "weights" of the tags.

Re: How Not to Sort by Average Rating (2009)

#129
post #54

Earlier quoted context omitted.

Half-Life fans are currently leaving negative reviews on Dota 2 because Valve won't make HL3. Recent reviews went from "overwhelmingly positive" to "mixed". https://arstechnica.com/gaming/2017/08/steam-reviewers-bomb-...

Valve should probably just farm out HL3 to Obsidian, and continue printing money with Steam.

Open-world Half-Life does sound intriguing.

Re: How Not to Sort by Average Rating (2009)

#130

Earlier quoted context omitted.

Perhaps the issue isn't the granularity of a single dimensional rating scale, but the lack of expressive options when in reality your feeling about something is complex and multifaceted. I've been really interested in the idea of emotive reviews as an alternative to single dimensional scores. The best idea I have at the moment is something akin to emoji reactions like you see on GitHub issues, finding a way to encode…

I envision a panel of emoticons akin to the Facebook reaction set, but where the user can select as many as they want to quickly convey different combinations of their reactions: (thumbs up) I liked this (heart) I loved this (thumbs down) I didn’t like this (smiling face) This made me happy or satisfied (frowning face) This made me sad or disappointed (surprised face) This made me surprised or impressed (angry face)…

How about vision based emotion recognition of viewers with cameras in the televisions and monitors? Sure sounds creepy and behaving different when observed etc. But I believe people will forget they are "observed" so the effect dimishs after a time. Than we would have a quite honest emotional feedback for movies. Even for specific scenes, for advertisment, etc
Post reply on HN