Live data from Hacker News

A popular self-driving car dataset is missing labels for hundreds of pedestrians

blog.roboflow.ai

61–70 of 202 posts

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#61
post #2

This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…

You would hope that the software would be resilient enough that missing or even faulty labels would not be the cause of problems.

It is by the way relatively easy to - once you reach some arbitrary threshold of precision - to detect such missing or wrong labels automatically, a simple suggestion UI with an accept / reject button could take care of supplying the bulk of the missing labels and correcting the bulk of the errors.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#62
post #12
post #6

Earlier quoted context omitted.

I understand your concern and share it myself. This is an important time and we should be really careful training these things. However, training as used in the real world isn't on a still frame only basis, it's used in sequence. And while a single frame might be missing a label, I bet that at-speed most everything important gets labeled correctly enough to be better than a distracted human driver, or the average hum…

> And while a single frame might be missing a label, I bet that at-speed most everything important gets labeled correctly enough to be better than a distracted human driver, or the average human driver for that matter. Would you bet a family member? That a distracted driver is a hazard does not mean other drivers are safe, or even saf er .

You currently do every day when you let your family members drive with senior citizens.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#63
post #11

It is a self-correcting problem: these pedestrains won’t be present in the next dataset.

People should learn not to go outside if they're not labelled.

Suddenly, being labelled doesn't sound so bad.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#64

Earlier quoted context omitted.

It does not make use aware of a new class of errors. Labeling issues is nothing new, but plenty of systems trained on them continue to work just fine. This is FUD.

Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling? Is there any reason - other than the combination of pure circumstance and gut feeling of the Data Scientist in charge of saying that it's good enough to deploy - that ~33% insanity in training doesn't become ~33% insanity in the system?

> Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling

Yes, it's called statistics and probability theory.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#65

Earlier quoted context omitted.

I'm not sure what point you're trying to make, but it has nothing to do with some mislabeled examples. Every system using supervised training assets has labeling issues, and the world hasn't ended yet. Take a look at Google Translate. This is only "scary" if you're ignorant of the problem. The OP is selling something, and it's in their best interest to spread FUD to sell it.

Nobody gets killed by Google Translate etc. messing up, so no sweat, throw together something that works 99% or the time. But for safety critical systems it's six nines reliability or GTFO.

Again, how does that have anything to do with a public data set? The amount of pure BS FUD in this topic is astounding.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#66

Earlier quoted context omitted.

Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling? Is there any reason - other than the combination of pure circumstance and gut feeling of the Data Scientist in charge of saying that it's good enough to deploy - that ~33% insanity in training doesn't become ~33% insanity in the system?

> Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling Yes, it's called statistics and probability theory.

That's correct. I know what goes in and what comes out, not what happens in the middle. How does ~33% insanity in become Edit: Parent was edited, was previously (paraphrased)

> I'm guessing you have no technical understanding of how this works

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#67

I work in the AV space. There's a lot of ambiguity in labeling. How should a crowd of people be annotated? A line of parked cars? A photograph of a car? Examine the training set!

A big thing is not that the example is missing, but that it counts as a negative example. I.e. if during training a ML system notices the ambiguous combination (i.e. a woman pushing a baby stroller or a crowd) and marks it as a pedestrian, then it gets penalized in a manner that teaches it to ignore these ambigious combinations and treat it as nothing; while in practice it should probably treat such ambiguous combina…

Right, what's really needed is a "clear road detector".

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#68

Earlier quoted context omitted.

It's scary when you consider two other factors: First, the AI hype train. People think that calling something "Artificial Intelligence" implies that it is artificial, yes, but also, critically, that it is intelligent. Many enthusiastic people, and also many policymakers, don't fully realize the extent to which machine learning is constrained by both the quality and nature of its training data, and the capabilities of…

I'm not sure what point you're trying to make, but it has nothing to do with some mislabeled examples. Every system using supervised training assets has labeling issues, and the world hasn't ended yet. Take a look at Google Translate. This is only "scary" if you're ignorant of the problem. The OP is selling something, and it's in their best interest to spread FUD to sell it.

That's probably the least charitable interpretation possible. No idea who the OP is, in any case.

It's scary that the only way we know how to build something that detects pedestrians in an image with any kind of reliability is to use training data.

You: All training data has labeling issues. Me: Training data is the only way we know how to build some aspects of systems. Other people here: Some of these systems are safety-critical.

I doesn't feel like FUD to be concerned that we need to have strategies for mitigating quality issues in training data. That if you don't have anything like that, then you cannot be in the market of building a safety-critical system. As you have identified, labeling issues are always there. Kind of how human error is always there. Okay, so now that you believe that, you have to find ways to mitigate it.

But when I hang around people who do self-driving cars, they don't have a good answer to how to mitigate this risk. Some of them don't even really believe in it (which is different than being ignorant of the problem), because they argue with enough data things will get good enough. It's all really sloppy.

Instead of just saying "This is FUD", why not tell us why you feel so reassured that this is not a problem? So far, you said "this always has been a problem and always will be a problem", "the world hasn't ended yet" and "Google translate". What?

re: The world hasn't ended yet: okay, but we also haven't been building self-driving cars

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#69

Earlier quoted context omitted.

Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling? Is there any reason - other than the combination of pure circumstance and gut feeling of the Data Scientist in charge of saying that it's good enough to deploy - that ~33% insanity in training doesn't become ~33% insanity in the system?

> Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling Yes, it's called statistics and probability theory.

> Yes it's called statistics and probability theory.

My understanding of statistics is:

- I can halve the % insanity by adding another 100% of good labels.

- If I want to reduce the insanity of labels to 1/33th of ~33% I need to add another 3200% of good labels.

- If I want to reduce the insanity to 0% I need to balance the bad labels with an infinite amount of good labels.

Is there anything I'm missing entirely except probability theory? Is probability theory the answer or is there something else?

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#70
post #28

Earlier quoted context omitted.

Given that nuclear power is significantly safer than other forms of power, are you asserting that the risks of self-driving cars are more about PR and perception than actual risk?

I'm saying that nuclear power could have been a safer energy option, but, in practice, the whole enterprise has been scuttled by a bunch of regrettably bad decisions that have pretty much destroyed everyone's trust. So now it doesn't really matter if it's safer, because it can no longer realistically be considered an option.

If everyone believes that everyone else will be irrational, then they themselves will not throw their support behind nuclear power, rendering their stance on nuclear de facto irrational. As Baby Boomers age, and Millennials/GenZ form a greater percentage of the voting population, we have an opportunity to press the reset button on nuclear. The younger generations don’t really have a solid opinion on the matter, and probably don’t think about it much.

Scientists are in the best position to influence the public and push for change. It is astonishing that scientific organizations have been timid about nuclear energy, or outright against it, because of a few people with the most extreme, empirically unjustified stance.

Post reply on HN