Live data from Hacker News

30% of Google's Emotions Dataset Is Mislabeled

surgehq.ai

21–30 of 146 posts

Re: 30% of Google's Emotions Dataset Is Mislabeled

#21

Earlier quoted context omitted.

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

you have experience with this, so you're probably the perfect person to ask this to- why didn't you just use the old (inaccurate) labels, perform some clustering based op and re-label then clusters? does that even make sense?

Doesn't this bias the labels?

Re: 30% of Google's Emotions Dataset Is Mislabeled

#22
post #21

Earlier quoted context omitted.

you have experience with this, so you're probably the perfect person to ask this to- why didn't you just use the old (inaccurate) labels, perform some clustering based op and re-label then clusters? does that even make sense?

Doesn't this bias the labels?

probably does, but idk how much.. if its mostly okay it's gonna be easier to correct it id assume

Re: 30% of Google's Emotions Dataset Is Mislabeled

#24
Ok, but now this person has done a bunch of unpaid work for them, just to publish an article, and now they can write some easy scripts to label any occurrences of 'daa+amn girl' as approval (etc. etc.) and in the end only 28% of the dataset will be mislabeled. The system works!

Re: 30% of Google's Emotions Dataset Is Mislabeled

#25
post #17
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

>Farm it out to low-paid non-native speakers The paper claims: >“All raters are native English speakers from India.”

https://www.heritagexperiential.org/language-policy-in-india...

>English, due to its ‘lingua franca’ status, is an aspiration language for most Indians – for learning English is viewed as a ticket to economic prosperity and social status. Thus almost all private schools in India are English medium. Many public schools, due to political compulsions, have the state’s official languages as the primary school language. English is introduced as a second language from grade 5 onwards.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#26

Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.

Maybe they already have a pool of workers in India and they are using it for all sort of tasks. If that's the case, they might have had to start a new process to get people in the US to label those sentences. Starting processes cost time and money and executive's political capital. Using an existing one is nearly for free.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#28
I have three questions now:

* How much (per comment) are these "native speakers from India" paid?

* How many comments do they have to label in an hour (or in a minute)? I guess it's more than 2 comments in a minute.

* What if the comment is sarcastic and this can only be understood from its context?

Re: 30% of Google's Emotions Dataset Is Mislabeled

#29

> let’s look at the labeling methodology described in the paper. To quote Section 3.3: > “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).” > “All raters are native English speakers from India.” This does not look good even on paper. No wonder the errors were abundant Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s

Isnt knowing the subreddit valuable information to determine sentiment?

Re: 30% of Google's Emotions Dataset Is Mislabeled

#30
post #27

> LETS FUCKING GOOOOO Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.

That is 100% enthusiasm, not ambiguous to me at all. But I'm also certain both my parents would read that as antagonistic.

It could also be 100% irony. Without context it's hard to tell.
Post reply on HN