Earlier quoted context omitted.
For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…
you have experience with this, so you're probably the perfect person to ask this to- why didn't you just use the old (inaccurate) labels, perform some clustering based op and re-label then clusters? does that even make sense?
30% of Google's Emotions Dataset Is Mislabeled
21–30 of 146 posts
Re: 30% of Google's Emotions Dataset Is Mislabeled
#22Earlier quoted context omitted.
you have experience with this, so you're probably the perfect person to ask this to- why didn't you just use the old (inaccurate) labels, perform some clustering based op and re-label then clusters? does that even make sense?
Doesn't this bias the labels?
Re: 30% of Google's Emotions Dataset Is Mislabeled
#23> LETS FUCKING GOOOOO Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#24Re: 30% of Google's Emotions Dataset Is Mislabeled
#25Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
>Farm it out to low-paid non-native speakers The paper claims: >“All raters are native English speakers from India.”
>English, due to its ‘lingua franca’ status, is an aspiration language for most Indians – for learning English is viewed as a ticket to economic prosperity and social status. Thus almost all private schools in India are English medium. Many public schools, due to political compulsions, have the state’s official languages as the primary school language. English is introduced as a second language from grade 5 onwards.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#26Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#27> LETS FUCKING GOOOOO Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
But I'm also certain both my parents would read that as antagonistic.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#28* How much (per comment) are these "native speakers from India" paid?
* How many comments do they have to label in an hour (or in a minute)? I guess it's more than 2 comments in a minute.
* What if the comment is sarcastic and this can only be understood from its context?
Re: 30% of Google's Emotions Dataset Is Mislabeled
#29> let’s look at the labeling methodology described in the paper. To quote Section 3.3: > “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).” > “All raters are native English speakers from India.” This does not look good even on paper. No wonder the errors were abundant Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s
Re: 30% of Google's Emotions Dataset Is Mislabeled
#30> LETS FUCKING GOOOOO Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
That is 100% enthusiasm, not ambiguous to me at all. But I'm also certain both my parents would read that as antagonistic.