Earlier quoted context omitted.
I was using MTurk for labeling about 10 years ago. To see the other side I also did a 1 Month stint as a MTurk worker, earning about $300. It is absolutely horrible work, and I used MTurk subreddit to find the "decent" jobs. I had the special firefox extension which ranked the job givers etc. All jobs were below 1st world minimum wage and were incredibly depressing. I think the adult content ones were the worst. "Jia…
I'm eager to see what nefarious things Google will do once they've mastered the art of identifying all the busses in a photo!
30% of Google's Emotions Dataset Is Mislabeled
141–146 of 146 posts
Re: 30% of Google's Emotions Dataset Is Mislabeled
#142Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
i mean humans arent great at reading emotions either. 30% error is probably human level
Re: 30% of Google's Emotions Dataset Is Mislabeled
#143Earlier quoted context omitted.
i mean humans arent great at reading emotions either. 30% error is probably human level
The film director Alfred Hitchcock once commented that in a tense scene, all he needed was a character showing a more or less neutral face, and viewers would read what he needed into it.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#144I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…
I think that in many of these cases the deciding factor is not only fluency in the language, but also the harder problem of context. Labelling individual sentences without context is hard enough, but what makes it worse is that it then spreads to the mistaken assumption that sentences can be analysed without context based on the initial training. I would argue that the very idea of "sentiment analysis" as applied to…
Oh. Shoot me now.
Note to future taggers: I am not suicidal.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#145Language is hard! Even I, a seasoned native internet dork, have trouble knowing if someone's comment is sarcasm, irony, or something in between. Also, new phrases emerge all the time that turn a phrase on its head, and it has a different emotion. How many feelings can you evoke with a simple, FUCK!
Re: 30% of Google's Emotions Dataset Is Mislabeled
#146Earlier quoted context omitted.
That's what I'm saying. Classifying one data point is very straightforward and brings negligible value to a company. Reliably classifying hundreds of thousands of them is very complicated and not at all easily supervised. And if your company's business model is based on applying trained models to $real_world_problem, it doesn't just bring a lot of value to your company, it's literally critical for its success, just l…
Are you trolling us here? The act of classifying is easy. It needs to be repeated many times. So you scale up by hiring many people to do it.
First of all, real-life data sets have hundreds of thousands of endpoints, and it's easy to classify a few hundred, maybe a few thousands endpoints for a single person. So scaling it up to a point where it's easy for every person involved in it requires hiring a team of dozens, or even 100+ people. That is absolutely not easy, especially not on a short term notice, and not when it's 100% a dead end job, so it's difficult to convince people to come do it in the first place. I'm yet to have met a single company whose idea of scaling it up involved something more elaborate than "we're gonna hire three freelancers". At 30k endpoints/person that's about an order of magnitude away from "easy", and transferring that order of magnitude to the hiring process ("we're gonna hire thirty freelancers") isn't trivial at all.
Second, it is absolutely not at all easy to supervise. QC for classification problems is comparable to QC for other easily-replicable but difficult to automate industrial processes, like semi-automatic manufacturing processes. There is ample literature on the topic, starting with the late period of the industrial revolution and going all the way to present times, and all of it suggests that it's a very hairy problem even without taking into account the human part. Perfect verification requires replicating the classification process. Verification by sampling makes it very hard to guarantee the accuracy requirements of the model. Checking accuracy post-factum poses the same problems.
This idea that classifying training models is a simple job that you can just outsource somewhere cheap is the framework of a bad strategy. Training data accuracy is absolutely critical. If you optimize for time and cost, you get exactly what you pay for: a rushed, cheap model.