Live data from Hacker News

30% of Google's Emotions Dataset Is Mislabeled

surgehq.ai

91–100 of 146 posts

Re: 30% of Google's Emotions Dataset Is Mislabeled

#91
At our university we do the clarification ourselves, usually with 3-5 classifiers per item. And it is surprising the high rate of disconformity between labelers (even in binary or 4-class classification). People in internet don't always understand sarcasm, so maybe we need to benchmark humans at this task. But yeah, this and similar datasets (in Spanish, for example) have tons of misclassified texts.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#92
post #38

Language is hard! Even I, a seasoned native internet dork, have trouble knowing if someone's comment is sarcasm, irony, or something in between. Also, new phrases emerge all the time that turn a phrase on its head, and it has a different emotion. How many feelings can you evoke with a simple, FUCK!

I feel like this is basically a subset of the translation problem, which you probably need actual artificial general intelligence for (because you need to be able to model another human mind to a certain degree). Here's a neat video covering the topic [1].

1: https://www.youtube.com/watch?v=GAgp7nXdkLU

Re: 30% of Google's Emotions Dataset Is Mislabeled

#93
post #80

Earlier quoted context omitted.

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

I was using MTurk for labeling about 10 years ago. To see the other side I also did a 1 Month stint as a MTurk worker, earning about $300. It is absolutely horrible work, and I used MTurk subreddit to find the "decent" jobs. I had the special firefox extension which ranked the job givers etc. All jobs were below 1st world minimum wage and were incredibly depressing. I think the adult content ones were the worst. "Jia…

I'm eager to see what nefarious things Google will do once they've mastered the art of identifying all the busses in a photo!

Re: 30% of Google's Emotions Dataset Is Mislabeled

#94

Wow, this explains a lot . I wonder if they're as inept when it comes to their search tech. The search result quality these days certainly speaks volumes.

Someone needs to write a sci-fi short story where in the future, a google AI is trained to maximize human happiness, but its ability to predict human happiness relies on this dataset of mislabeled human emotions farmed out to underpaid Indian workers.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#95
I can see how this issue will happen frequently, in diverse domains and abundantly.

In other words, the prediction is that the most likely outcome will be a lot of AI objects trained to be quite imbecile and will be optimal at that.

The danger is that real people might be assumed to be guilty of things due to AI trained and automated imbecility.

It's an ethical problem for the AI community and product designers.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#96

Earlier quoted context omitted.

Searching and labeling are vastly different areas. Google already proved their expertise in search years ago - what they do now is expand and adapt to changes.

When I was small, I decided that our neatly ordered little drawers of Lego would be much better if they were jumbled up - every drawer would then contain a sort of average collection so I would be able to just open one at random to get the part I needed. It seems that Google have applied that philosophy to search results.

Did you design Amazon's warehousing system?

Re: 30% of Google's Emotions Dataset Is Mislabeled

#98
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

Why can’t you do the dataset several times with different labellers, building up a more statistical probability for a label than a pure guarantee. It’s worth noting that different cultures have different interpretations of emotions too (famously Russians don’t smile much even though they are extremely helpful in my experience, I’m still not quite sure what the head wobble in India actually means let alone some of the finer misinterpretations that are going on when I lived in Japan).

Re: 30% of Google's Emotions Dataset Is Mislabeled

#99
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

> I’m still shocked how low the quality of Mechanical Turk was for just sentiment

Given that MT workers earn based on how quickly they complete a task, and not how accurate. I'm struggling to understand how anyone would expect quality.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#100

Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.

And this is even full-time in the sense of 24/7! If you assume that people only work 8 hours a day, it would take 22 people to do it within a month. Though, on second thought, with current inflation rates people probably need to work 720 hours a month ;P

It was supposed to be 8 hour days ... maybe I messed up the calculation. Either way, even if it's ten-fold: it's really not all that much money for a company Google's size, especially considering the "cheap" alternative gives poor results (which also translate to monetary costs).
Post reply on HN