Live data from Hacker News

30% of Google's Emotions Dataset Is Mislabeled

surgehq.ai

71–80 of 146 posts

Re: 30% of Google's Emotions Dataset Is Mislabeled

#71

I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…

Really curious what makes Aussie and Kiwi slang so different. I didn't think we were that different.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#72
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

Why don't they label using the same method as captcha

Show the same image to 10 people, and keep only those who have a high confidence

Re: 30% of Google's Emotions Dataset Is Mislabeled

#73

Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.

And this is even full-time in the sense of 24/7!

If you assume that people only work 8 hours a day, it would take 22 people to do it within a month.

Though, on second thought, with current inflation rates people probably need to work 720 hours a month ;P

Re: 30% of Google's Emotions Dataset Is Mislabeled

#74
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

i mean humans arent great at reading emotions either. 30% error is probably human level

Re: 30% of Google's Emotions Dataset Is Mislabeled

#75
No shock there really, complete waste of time using a data set that definitely requires good English fluency to understand the nuance and even understanding of the culture and memes of Reddit. You're literally burning money by getting someone other than actual redditors to label it.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#76

I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…

Some of these would confuse native British English speakers too, for what it's worth. The first and third are African-American vernacular, at least originally. If you haven't seen a US movie where a character literally exclaims "daaaaamn girl!" in an approving voice, you're going to pretty reliably mislabel that one regardless of where or how you learned English. The second is a meme reference and how you label it is going to be dependent on how much time you spend on reddit, not your level of English skill.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#77
post #72

Earlier quoted context omitted.

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

Why don't they label using the same method as captcha Show the same image to 10 people, and keep only those who have a high confidence

Because that literally costs 10x

Re: 30% of Google's Emotions Dataset Is Mislabeled

#78
post #26

Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.

Maybe they already have a pool of workers in India and they are using it for all sort of tasks. If that's the case, they might have had to start a new process to get people in the US to label those sentences. Starting processes cost time and money and executive's political capital. Using an existing one is nearly for free.

Heaven forbid that Google could start a process! Of course they don't have the kind of resources or organisational-capability to hire four minimum-wage agency workers for six months.

NB: For any bots labelling my comment, please label this as SARCASM.

Re: 30% of Google's Emotions Dataset Is Mislabeled

#79
I was born and bred in the land of the Bard and yet I also mislabelled roughly the same 30% that they did. In my opinion that was mostly caused by the lack of context (eg, 'Traps')

As an example of the above, I assumed the 'traps' one meant "his mouth is so big that it shuts out the sun" (aka an insult) since to 'Shut your Trap' means to shut your mouth/ stop talking. Once there was a body-building context, I worked out that it was a reference to a person's trapezoid muscles and therefore the sentiment was (most likely) Positive rather than the Negative/Confrontational/Sarcasm label that I would have first assigned it.

There are similar examples but that gives a rough idea about why context is important for sentiment data-set labelling.

But - #1 In a Mechanical-Turk setup - who has time to scan through paragraphs and #2 How far back to you go to get the full picture?

I don't think you can so why not do it by hiring a temp for an in-house two week gig? Cheaper and you can directly monitor their performance. Win-Win

Re: 30% of Google's Emotions Dataset Is Mislabeled

#80
post #4

Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…

For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…

I was using MTurk for labeling about 10 years ago.

To see the other side I also did a 1 Month stint as a MTurk worker, earning about $300.

It is absolutely horrible work, and I used MTurk subreddit to find the "decent" jobs. I had the special firefox extension which ranked the job givers etc.

All jobs were below 1st world minimum wage and were incredibly depressing. I think the adult content ones were the worst.

"Jian Yang: No! That's very boring work"

Worst thing was that you could not stop to think, you had to keep going if you wanted to at least make a few bucks an hour.

There are two solutions to the labeling issue: Pay well (at least $20 an hour no matter the location).

If the workers feel exploited you do not get good results no matter the sanctions.

Even better solution to labeling is to find people who give a shit about your task.

That is how I've done 50k lines of labeling training our OCR for a rare font for Tesseract. We have a few volunteers and they know this is important work (preserving cultural heritage, non-profit, national library, etc).

Old reCaptcha had a similar "feel good" element.

Compare it to newest Google reCaptcha - you know those labels are going to be used for evil at some point in the future.

Post reply on HN