Live data from Hacker News

A popular self-driving car dataset is missing labels for hundreds of pedestrians

blog.roboflow.ai

171–180 of 202 posts

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#171
post #57

I acknowledge the issues in the dataset and that it has a lot of stars on github because it's from Udacity; but calling it 'a popular self-driving car dataset' is misleading as it implies this dataset is popularly used for self-driving cars when it is in fact only a small dataset Udacity uses to teach the basics of training neural networks for self-driving cars. I've been involved in the autonomous vehicle industry f…

Are these larger datasets routinely subject to the same kind of inspection this titanic.csv of self-driving car datasets?

I hope so. I've personally tested Scale's labeling service and it was much higher quality than this dataset. But it's a pretty secretive industry so I'd bet some companies' data is better than others.

It'd be interesting if the NHTSB had a held-back "test set" they used to evaluate self driving cars before letting them on the road.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#172

Isn't training against subset of data and then validating against rest a common practice? It wouldn't detect all the mislabeling but should detect some indicating that manual inspection is required, assuming error isn't very systematic.

It is, and there are some interesting techniques published recently to help mitigate things like this. But if you don't have a good ground truth you're at the very least flying blind and at worst feeding garbage in and getting garbage out; your models will learn what you tell them to learn.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#173
post #18

These are manually tagged stills, right? Not video? That's a data set for training CAPCHA breakers, not self-driving. You need to use video, where you get to see the same objects at different ranges. Recognition gets better as you get closer. Then track the recognized objects backwards to when they first appear, and try to recognize them at smaller sizes.

I've never built a self driving car but if I were to give it a go, identifying what's in a frame would be the first step I'd tackle. From there you'd add more layers to the stack to get a full understanding of the world, predict what will happen next temporally, and then choose which actions to take.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#174

Earlier quoted context omitted.

> I assume that special care will be taken and the cars themselves will be quite risk-averse. Bold assumption that's already got a counterpoint: https://en.wikipedia.org/wiki/Death_of_Elaine_Herzberg

Wikipedia already has a list: https://en.wikipedia.org/wiki/List_of_self-driving_car_fatal... Along with your pedestrian death there were five (!!!) driver deaths. There's a reason why places like Germany don't want Tesla to use the term 'autopilot' because it's reckless endangerment.

>There's a reason why places like Germany don't want Tesla to use the term 'autopilot' because it's reckless endangerment.

If Tesla weren't a US company I'm sure they'd have already been sued for billions for labeling their assisted driving system autopilot.

The US is the country where you can pay millions for not having a sticker on the microwave showing that you shouldn't dry your cat in there.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#175

My knee-jerk is to be upset because the risk factor of a poorly functioning pedestrian classifier is obviously higher than something like a sentiment analyzer, but at the same time, this dataset is just for educational purposes, right? Is Udacity actively recommending people use this in production settings?

They claim[1] they're working on building an "open source self driving car" but it looks like the project hasn't had much activity recently.

[1] https://github.com/udacity/self-driving-car

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#176
post #5

What's the dataset size? If there's billions of pedestrians tagged and a couple hundreds are missing, would it actually have a big impact on the training? Also looks like a lot of the people in those images are off the road. What's the current standard in AV? Is everything tagged on the sidewalk? (Genuine questions)

The dataset is 15,000 images. Breakdown of the number of labels per class (post fixes) is here: https://i.imgur.com/bOFkueI.png

Not all of the pedestrians and cyclists were on the sidewalk, no (eg the kid on his bike in the road and the lady with a stroller in a crosswalk).

I stuck with what it looked like the conventions of the original dataset were (all people labeled as pedestrians whether on the road or not). They just didn't do it very well or consistently.

I do also think that makes the most sense in this context; if you were building a self driving car this layer of the stack would want to know where the people are; higher up you can combine that with where you know the roads/crosswalks/stoplights, etc (and the delta of their position between frames) to make predictions about where they might go next so your car can act accordingly.

For example, a stationary pedestrian at a corner will probably cross the street when the light turns green; if you're turning you need to factor that in.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#177

This seems like the kind of problem that is suitable for and important enough to throw hundreds or thousands of (volunteer?) people at. Imagine assembling a massive public corpus of human verified training data - we'd collectively be that much closer to a technology which will change society! This problem is screaming for a consortium effort on behalf of major corporate entities in the space, who could each benefit w…

It's interesting that it was a public dataset (ImageNet) that kickstarted the amazing strides in computer vision over the last decade! More public data would be awesome (if high quality).

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#178

Earlier quoted context omitted.

What if its very much better average accident rates? This isn't black-and-white.

No, it certainly isn't black and white. Indeed 'much better' is hard to even define when human drivers cover an enormous amount of miles per accident, miles driven are heterogenous in terms of risk, there isn't even necessarily a universally accepted classification of accident severity or whether drivers should be excluded from the sample as being 'at fault' to an unacceptable degree. Plus the AV software isn't stayi…

Sophistry. 'Much better' can be very clear, in terms of death or injury, or property damage, or insurance claims, or half a dozen reasonable measures.

Sure it takes miles to determine what's better. Once automated driving is happening in millions (instead of hundreds) of cars on the road, it will take only days to measure.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#179
post #119

Earlier quoted context omitted.

IMO to be "scared" of self driving cars they need to be more dangerous than any other random car on the street today, and the bar for that is pretty low.

Nope, with other cars you know the factors that increase your risks (fog, drunk driving, distractions). You can make decisions, like not getting into a car with a drunk friend. With machine learning you never know when it might mistake the back of a semi for an overpass or soemthing.

No pedestrian makes the decision to get hit by a drunk driver, which is the more relevant situation in the context of this post I think

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#180

Earlier quoted context omitted.

No, it certainly isn't black and white. Indeed 'much better' is hard to even define when human drivers cover an enormous amount of miles per accident, miles driven are heterogenous in terms of risk, there isn't even necessarily a universally accepted classification of accident severity or whether drivers should be excluded from the sample as being 'at fault' to an unacceptable degree. Plus the AV software isn't stayi…

Sophistry. 'Much better' can be very clear, in terms of death or injury, or property damage, or insurance claims, or half a dozen reasonable measures. Sure it takes miles to determine what's better. Once automated driving is happening in millions (instead of hundreds) of cars on the road, it will take only days to measure.

I mean, the 'half a dozen reasonable measures' is a problem, not a solution, when they're not all saying the same thing. And sure, it only takes days before we know the latest version of the software actually isn't safer than the average human. And a lot of unnecessary deaths, and the likelihood the fix will cause other unnecessary deaths instead [maybe more, maybe less]. It's frankly sociopathic to dismiss the possibility this might be a problem as sophistry.
Post reply on HN