30% of Google's Emotions Dataset Is Mislabeled
1–10 of 146 posts
Re: 30% of Google's Emotions Dataset Is Mislabeled
#2Re: 30% of Google's Emotions Dataset Is Mislabeled
#3Since words only have meaning within a context, your model should reflect that somehow.
What wasn't really explored I'm this article was to what quantitative degree context sensitivity matters. The counter examples are great, but how can we measure the relationship between amount of context and labelling accuracy?
Re: 30% of Google's Emotions Dataset Is Mislabeled
#4Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something you outsource (or if you do, you need an additional in-house validation pipeline to identify dirty labels).
Re: 30% of Google's Emotions Dataset Is Mislabeled
#5Wow, this explains a lot . I wonder if they're as inept when it comes to their search tech. The search result quality these days certainly speaks volumes.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#6this is a genuinely great read. Author does a great job providing examples where context is critical, and explains how the dataset not only has labeling errors, an even deeper problem is how it models language in general. Since words only have meaning within a context, your model should reflect that somehow. What wasn't really explored I'm this article was to what quantitative degree context sensitivity matters. The…
My dog has a better understanding of context than any AI I've ever met. (I mean that sincerely - since becoming a pet owner it's something I've really marveled at).
Re: 30% of Google's Emotions Dataset Is Mislabeled
#7Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
I wonder how others do this kind of thing. Assuming I have 100k text blurbs I want to classify of sentiment and/or emotion and want at least 95% accuracy, what are the options?
I’m still shocked how low the quality of Mechanical Turk was for just sentiment (positive/negative/neutral/unsure), 99% of the classifications were just random. We narrowed our section for workers to higher-qualified ones, for that matter.
What a giant waste of money and time that was, because supposedly it’s the canonical use case for it.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#8Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
Funnily enough, though, many ML engineers and data scientists I know (even those at Google, etc., who depend on human-annotated datasest) aren't familiar with these kinds of errors. At least in my experience, many people rarely inspect their datasets -- they run their black box ML pipelines and compute their confusion matrices, but rarely look at their false positive/negatives to understand more viscerally where and why their models might be failing.
Or when they do see labeling errors, many people chalk it up to "oh, it's just because emotions are subjective, overall I'm sure the labels are fine" without realizing the extent of the problem, or realizing that it's fixable and their data could actually be so much better.
One of my biggest frustrations actually is when great engineers do notice the errors and care, and try to fix them by improving guidelines -- but often the problem isn't the guidelines themselves (in this case, for example, it's not like people don't know what JOY and ANGER are! creating 30 pages of guidelines isn't going to help), but rather that the labeling infrastructure is broken or nonexistent from the beginning. Hence why Surge AI exists, and we're building what we're building :)
Re: 30% of Google's Emotions Dataset Is Mislabeled
#9Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#10> “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).”
> “All raters are native English speakers from India.”
This does not look good even on paper. No wonder the errors were abundant
Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s