This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…
A popular self-driving car dataset is missing labels for hundreds of pedestrians
71–80 of 202 posts
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#72Only when the machines start to complain when they are fed shitty data, we can talk about them being fit to drive.
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#73Earlier quoted context omitted.
>Would you bet a family member? "think of the children!" Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'. Thousands of people are dying every day in cars. This is not something we need to wait for it to be perfect. It only need…
> Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'. Maybe logically that makes sense but from an ethical perspective I argue it's much more complicated than that (e.g. the trolley problem) In the current system if a human is at…
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#74This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…
> This is really scary. No, it's not even remotely "really scary". No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space or is represented of the impact a few mislabelled samples have on the ability of these systems & algorithms to generalize.
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#75I acknowledge the issues in the dataset and that it has a lot of stars on github because it's from Udacity; but calling it 'a popular self-driving car dataset' is misleading as it implies this dataset is popularly used for self-driving cars when it is in fact only a small dataset Udacity uses to teach the basics of training neural networks for self-driving cars. I've been involved in the autonomous vehicle industry f…
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#76Earlier quoted context omitted.
> Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'. Maybe logically that makes sense but from an ethical perspective I argue it's much more complicated than that (e.g. the trolley problem) In the current system if a human is at…
I'm voting for the "less dead people" option. Mostly because I'm a selfish person, been in automobile accidents caused by lapsing human attention, and I want it to be less likely that I'll die in a car crash.
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#77Earlier quoted context omitted.
> Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling Yes, it's called statistics and probability theory.
That's correct. I know what goes in and what comes out, not what happens in the middle. How does ~33% insanity in become Edit: Parent was edited, was previously (paraphrased) > I'm guessing you have no technical understanding of how this works
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#78Earlier quoted context omitted.
> Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling Yes, it's called statistics and probability theory.
> Yes it's called statistics and probability theory. My understanding of statistics is: - I can halve the % insanity by adding another 100% of good labels. - If I want to reduce the insanity of labels to 1/33th of ~33% I need to add another 3200% of good labels. - If I want to reduce the insanity to 0% I need to balance the bad labels with an infinite amount of good labels. Is there anything I'm missing entirely exce…
Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians
#79Earlier quoted context omitted.
People should learn not to go outside if they're not labelled.
What a funny future it'd be, if we have to wear something distinctive (giant QR codes?) so we don't get killed outside. As a side effect it would make tracking us much easier...