Live data from Hacker News

We in-housed our data labelling

ericbutton.co

11–20 of 51 posts

Re: We in-housed our data labelling

#11
post #9
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

Anyone have the link to that thread?

The posts last week about the factory worker monitoring startup? There were at least 3 posts (and I think dang let the dupes through, since some mentioned YC, and HN moderates less in such cases):

https://news.ycombinator.com/item?id=43175023

https://news.ycombinator.com/item?id=43170850

https://news.ycombinator.com/item?id=43180133

Re: We in-housed our data labelling

#12
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

Some of the big-dollar contracts $employer has involve financial penalties if performance metrics aren't up to standard.

Re: We in-housed our data labelling

#13
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

Some of the big-dollar contracts $employer has involve financial penalties if performance metrics aren't up to standard.

At the same time, many organisations getting work done through platforms like Mechanical Turk set their piece rate to make sure all but the worst workers will make at least minimum wage.

Re: We in-housed our data labelling

#14
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

What do you mean “didn’t age well”, its a brand new article. It hasn’t aged at all.

Re: We in-housed our data labelling

#15
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

> But rather, under a punishing set of Kafkaesque rules, like someone was thinking only of computer programs, oops. "Gamified", with huge negative points penalties and everything. To be under threat of not getting paid at all.

I'm not defending these practices, but to share some context:

One of the problems with getting workers to review ML output is it's incredibly, unbelievably boring. When the task is to review model output you're going to hit the 'approve' button 99% of the time - and when you're being paid for speed, nothing's faster than hitting the approve button.

So understandably a decent number of folks will just zone out, maybe put youtube on in another window, and sit there hitting approve 100% of the time. That's just human nature when dealing with such an incredibly dull task - I know I don't pay attention when I have to do my annual refresher training on how to sit in a chair.

This sort of thing is a big problem for things like airport baggage scanner operators; pilots with their planes on autopilot; lifeguards; casino CCTV operators; and suchlike. There are loads of studies about this kind of stuff.

This makes getting good quality ML output reviews quite tricky. There are ways to do it, though, and you don't have to resort to negative income!

Re: We in-housed our data labelling

#16
> Still, expert reviewers will occasionally disagree in their labelling. To ensure quality, an audio clip [box characters], at which point [...]

Have they censored their own article?

Re: We in-housed our data labelling

#18
> Failing a test will cost a user 600 points, or roughly the equivalent of 15 minutes of work on the platform. A correctly tuned penalty system removes the need for setting reviewer accuracy minimums; poor performers will simply not earn enough money to continue on the platform.

This still sets a reviewer accuracy minimum, but it is determined implicitly by the arbitrary test penalty instead of consciously chosen based on application requirements. I don't see how that's an improvement. If you absolutely want to have negative earnings, it would make more sense to choose a reviewer accuracy minimum to aim for, and then determine the penalty that would achieve that target, instead of the other way around.

Moreover, a reviewer earning nothing on expectation under this scheme (they work for 15 minutes, then fail a test, and have all their earnings wiped out) could team up with a second reviewer with the same problem, submitting their answer only when both agree, and as long as their errors aren't 100% correlated, they would end up with positive expected earnings they could split between them.

This clearly indicates that the incentive scheme as designed doesn't capture the full economic value of even lower-quality data when processed appropriately. Of course you can't expect random reviewers to spontaneously work together in this way, so it's up to the data consumer to combine the work of multiple reviewers as appropriate.

Trying to get reliable results from humans by exclusively hiring the most reliable ones can only get you so far; you can do much better by designing systems to use redundancy to correct errors when they inevitably do appear. Ironically, this is a case where treating humans as fallible cogs in a big machine would be more respectful.

Re: We in-housed our data labelling

#19
post #6

I think this didn't age well, for HN, and it prompts some serious questions about our techbro startup culture. > Obvious but necessary: to incentivize productive work, we tie compensation to the number of characters transcribed, and assess financial penalties for failed tests (more on tests below). Penalties are priced such that subpar performance will result in little to no earnings for the labeller. So, these aren'…

> To be under threat of not getting paid at all.

If they are operating as described, it’s almost certainly illegal. They deserve to be hit with a nice, fat PAGA lawsuit. These workers would have to satisfy the “ABC test” to be exempt from minimum wage obligations, and it’s a difficult standard to meet: https://www.labor.ca.gov/employmentstatus/abctest/

> I want to ask who is advising these startups regarding how they think of their place in the world, relative to other humans?

To me, this has been one of the most dispiriting things to witness in the last few years: not just the normalization, but the outright glorification, of indecency. Shameful.

Re: We in-housed our data labelling

#20
> All labellers are either licensed pilots or controllers (or VATSIM pilots/controllers).

I would think such people can make better money by actually working as a pilot or controller?

Post reply on HN