I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…
30% of Google's Emotions Dataset Is Mislabeled
71–80 of 146 posts
Re: 30% of Google's Emotions Dataset Is Mislabeled
#72Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…
Show the same image to 10 people, and keep only those who have a high confidence
Re: 30% of Google's Emotions Dataset Is Mislabeled
#73Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.
If you assume that people only work 8 hours a day, it would take 22 people to do it within a month.
Though, on second thought, with current inflation rates people probably need to work 720 hours a month ;P
Re: 30% of Google's Emotions Dataset Is Mislabeled
#74Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
Re: 30% of Google's Emotions Dataset Is Mislabeled
#75Re: 30% of Google's Emotions Dataset Is Mislabeled
#76I assumed I was quite fluent in English, even in slangs, having seen a fair share of both American and British movies. Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them. Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional s…
Re: 30% of Google's Emotions Dataset Is Mislabeled
#77Earlier quoted context omitted.
For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…
Why don't they label using the same method as captcha Show the same image to 10 people, and keep only those who have a high confidence
Re: 30% of Google's Emotions Dataset Is Mislabeled
#78Let's say you can label 2 comments a minute, you'd have to spend 3,625 work-hours to label comments, or about five people working full-time for a month. How much money did they save by using cheaper labour from India? Basically bugger all, and the money is wasted, too. Penny wise, pound foolish.
Maybe they already have a pool of workers in India and they are using it for all sort of tasks. If that's the case, they might have had to start a new process to get people in the US to label those sentences. Starting processes cost time and money and executive's political capital. Using an existing one is nearly for free.
NB: For any bots labelling my comment, please label this as SARCASM.
Re: 30% of Google's Emotions Dataset Is Mislabeled
#79As an example of the above, I assumed the 'traps' one meant "his mouth is so big that it shuts out the sun" (aka an insult) since to 'Shut your Trap' means to shut your mouth/ stop talking. Once there was a body-building context, I worked out that it was a reference to a person's trapezoid muscles and therefore the sentiment was (most likely) Positive rather than the Negative/Confrontational/Sarcasm label that I would have first assigned it.
There are similar examples but that gives a rough idea about why context is important for sentiment data-set labelling.
But - #1 In a Mechanical-Turk setup - who has time to scan through paragraphs and #2 How far back to you go to get the full picture?
I don't think you can so why not do it by hiring a temp for an in-house two week gig? Cheaper and you can directly monitor their performance. Win-Win
Re: 30% of Google's Emotions Dataset Is Mislabeled
#80Anyone who's dealt with any kind of human-annotated datasets would be familiar with these kind of errors. It's hard enough to get good clean labels from motivated, native-English speaking annotators. Farm it out to low-paid non-native speakers, and these kind of issues are inevitable. Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something yo…
For a sentiment / emotion classification project, we (2 founders) just ended up doing most of the labeling ourselves. It was a big grind, but given how abysmal the performance of “crowd-sourced” solutions are (eg Amazon Mechanical Turk), and how incredibly important the quality of these labels are for training a model, it made the most sense. I wonder how others do this kind of thing. Assuming I have 100k text blurbs…
To see the other side I also did a 1 Month stint as a MTurk worker, earning about $300.
It is absolutely horrible work, and I used MTurk subreddit to find the "decent" jobs. I had the special firefox extension which ranked the job givers etc.
All jobs were below 1st world minimum wage and were incredibly depressing. I think the adult content ones were the worst.
"Jian Yang: No! That's very boring work"
Worst thing was that you could not stop to think, you had to keep going if you wanted to at least make a few bucks an hour.
There are two solutions to the labeling issue: Pay well (at least $20 an hour no matter the location).
If the workers feel exploited you do not get good results no matter the sanctions.
Even better solution to labeling is to find people who give a shit about your task.
That is how I've done 50k lines of labeling training our OCR for a rare font for Tesseract. We have a few volunteers and they know this is important work (preserving cultural heritage, non-profit, national library, etc).
Old reCaptcha had a similar "feel good" element.
Compare it to newest Google reCaptcha - you know those labels are going to be used for evil at some point in the future.