Live data from Hacker News

30% of Google's Emotions Dataset Is Mislabeled

surgehq.ai

61–70 of 146 posts

Re: 30% of Google's Emotions Dataset Is Mislabeled

#61
post #56

Earlier quoted context omitted.

> I’m still shocked how low the quality of Mechanical Turk was I don't know about Mechanical Turk, but there is a crowdourcing platform by Yandex. The pay is so low that the reasonable way to earn something is to find a task that is not properly validated and put random answers there from multiple accounts (because there are speed limits). Usually those are tasks by naive foreign companies not knowing about validatio…

Why shouldn’t someone do it for five dollars per hour, and do it properly lest they get fired? Seems like a very easy job and easy to supervise.

Because they can do more rolling carts around in Wallmart.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#62
I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies.

Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them.

Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional slangs of which I cannot understand all the nuances.

A few examples from this dataset, that I would not have labeled correctly:

- daaaaaamn girl! – mislabeled as ANGER

- [NAME] wept. – mislabeled as SADNESS

- [NAME] is bae, how dare you. – mislabeled as ANGER

And don't get me started on Australian/NZ slang. It's a completely different world.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#63

I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…

I think that in many of these cases the deciding factor is not only fluency in the language, but also the harder problem of context. Labelling individual sentences without context is hard enough, but what makes it worse is that it then spreads to the mistaken assumption that sentences can be analysed without context based on the initial training.

I would argue that the very idea of "sentiment analysis" as applied to individual tweets is flawed... and that's even before we get into the much, much harder problems of sarcasm and irony.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#64
I work at a company that focuses on automating the data labelling process for computer vision. It is clear that generating massive amounts of labels, either by hand or automatically, without the ability to ensure a consistent level of quality across the dataset, is a problem. Which is why we are investing in automating the QA process for training data so mistakes like these don't happen:

https://blog.encord.com/post/automating-the-assessment-of-tr...

Re: 30% of Google's Emotions Dataset Is Mislabeled

#65
post #64

I work at a company that focuses on automating the data labelling process for computer vision. It is clear that generating massive amounts of labels, either by hand or automatically, without the ability to ensure a consistent level of quality across the dataset, is a problem. Which is why we are investing in automating the QA process for training data so mistakes like these don't happen: https://blog.encord.com/post/…

That's interesting, I took the opposite approach when realizing how bad label quality can be and built a site where everything is manually checked by moderators.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#66
post #38

Language is hard! Even I, a seasoned native internet dork, have trouble knowing if someone's comment is sarcasm, irony, or something in between. Also, new phrases emerge all the time that turn a phrase on its head, and it has a different emotion. How many feelings can you evoke with a simple, FUCK!

And this is probably going to get even worse the more automatic classification is used to promote or silence content. A pretty interesting result of this is what I'd call "TikTok speak", where words are replaced, either by similar sounding ones ("porn" => "corn", often times just the corn emoji) or by neologisms ("to kill" => "to unalive"), in the hope of getting around the filters. This turns natural language on the…

The core of the problem is unsolvable, since any automated system can be defeated by a sufficiently motivated human.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#67
post #29

> let’s look at the labeling methodology described in the paper. To quote Section 3.3: > “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).” > “All raters are native English speakers from India.” This does not look good even on paper. No wonder the errors were abundant Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s

Isnt knowing the subreddit valuable information to determine sentiment?

But it also applies to the model. If you label data taking into account the context (which subreddit this was posted), your model also must take this into account, or it will be wrong as well. If the same sentence can be labeled two different ways depending on the source, then, the model must also know the source or it won't know what to do. But then you didn't create a general language model, you created a reddit language model.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#68
post #56

Earlier quoted context omitted.

> I’m still shocked how low the quality of Mechanical Turk was I don't know about Mechanical Turk, but there is a crowdourcing platform by Yandex. The pay is so low that the reasonable way to earn something is to find a task that is not properly validated and put random answers there from multiple accounts (because there are speed limits). Usually those are tasks by naive foreign companies not knowing about validatio…

Why shouldn’t someone do it for five dollars per hour, and do it properly lest they get fired? Seems like a very easy job and easy to supervise.

Money isn't what gets people motivated most of the time. This job is boring and unfulfilling, you'll get poor results unless you're offering a life changing amount of money.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#69
post #56

Earlier quoted context omitted.

Why shouldn’t someone do it for five dollars per hour, and do it properly lest they get fired? Seems like a very easy job and easy to supervise.

Because they can do more rolling carts around in Wallmart.

I mean for places where $5 per hour is an attractive wage.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#70
post #59
post #56

Earlier quoted context omitted.

Why shouldn’t someone do it for five dollars per hour, and do it properly lest they get fired? Seems like a very easy job and easy to supervise.

If it's such an easy job, why outsource it instead of doing it yourself?

Because there’s other stuff you need to do that you can’t outsource.
Post reply on HN