Live data from Hacker News

A popular self-driving car dataset is missing labels for hundreds of pedestrians

blog.roboflow.ai

31–40 of 202 posts

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#31
post #28

Earlier quoted context omitted.

It's scary when you consider two other factors: First, the AI hype train. People think that calling something "Artificial Intelligence" implies that it is artificial, yes, but also, critically, that it is intelligent. Many enthusiastic people, and also many policymakers, don't fully realize the extent to which machine learning is constrained by both the quality and nature of its training data, and the capabilities of…

Given that nuclear power is significantly safer than other forms of power, are you asserting that the risks of self-driving cars are more about PR and perception than actual risk?

I'm saying that nuclear power could have been a safer energy option, but, in practice, the whole enterprise has been scuttled by a bunch of regrettably bad decisions that have pretty much destroyed everyone's trust. So now it doesn't really matter if it's safer, because it can no longer realistically be considered an option.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#32
post #28

Earlier quoted context omitted.

It's scary when you consider two other factors: First, the AI hype train. People think that calling something "Artificial Intelligence" implies that it is artificial, yes, but also, critically, that it is intelligent. Many enthusiastic people, and also many policymakers, don't fully realize the extent to which machine learning is constrained by both the quality and nature of its training data, and the capabilities of…

Given that nuclear power is significantly safer than other forms of power, are you asserting that the risks of self-driving cars are more about PR and perception than actual risk?

I think the point is rather that it's an unpredictable-stakes game. That is, if you make few enough mistakes, you'll be fine for a long time until you make the wrong mistake at the wrong time and it kills somebody.

Maybe a better analogy is driving a car, that is, nuclear power and self-driving cars are both vaguely like driving a car.

However, my intuition is that nuclear power has less variables than driving a car. I say intuition because I don't know much about neither.

I think the point is that self-driving cars trained on poor data are no better than poorly trained superhumans at driving.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#33
post #13

Well, if an autonomous vehicle outfit were running unmonitored Level 4 vehicles on public roads using only an open source data set I'd be worried. Even if it was labelled thoroughly and correctly, there isn't nearly enough data in any open source dataset to train an autonomous vehicle perception system that can operate safely without human supervision. This is not a safety critical issue.

The Uber vehicle that ended up killing the pedestrian in Arizona was running under conditions similar to the ones that you outline here. The only notable difference was that a person was monitoring the vehicle.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#34
post #17
post #9

Earlier quoted context omitted.

A line of parked cars absolutely needs to be labeled as individual cars. Any one of them could pull out in front of you at any moment.

Indeed, whilst additionally presenting the chance of a door opening or an obscured pedestrian stepping out from between them.

Yup. You need to consider not only what you can see, but what you can't see. And this is a harder problem.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#35
post #11

It is a self-correcting problem: these pedestrains won’t be present in the next dataset.

People should learn not to go outside if they're not labelled.

> People should learn not to go outside if they're not labelled.

I worry what will happen when this idea breeds with the "why worry about privacy if you've got nothing to hide?" fallacy.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#36
post #12

Earlier quoted context omitted.

> And while a single frame might be missing a label, I bet that at-speed most everything important gets labeled correctly enough to be better than a distracted human driver, or the average human driver for that matter. Would you bet a family member? That a distracted driver is a hazard does not mean other drivers are safe, or even saf er .

>Would you bet a family member? "think of the children!" Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'. Thousands of people are dying every day in cars. This is not something we need to wait for it to be perfect. It only need…

There's still some nuance that's important. If self-driving cars always sacrifice other road users to protect the driver, self driving cars could be reduce death/injury overall, but there's a question about whether or not this behavior is ethical. And if the training datasets consistently label cars but not other road users, then this bias could be baked in completely by accident.

It only needs to be better.

A Pedestrian likely has a different definition of "better" than the car driver.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#37
post #6
post #2

This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…

I understand your concern and share it myself. This is an important time and we should be really careful training these things. However, training as used in the real world isn't on a still frame only basis, it's used in sequence. And while a single frame might be missing a label, I bet that at-speed most everything important gets labeled correctly enough to be better than a distracted human driver, or the average hum…

Your intuition should be that almost everything is incorrectly labelled. And even more so on a large scale sequential basis.

If you look at public datasets, it's more often that things are incorrectly tagged/labeled rather than correctly.

Entropy is a real thing.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#38
post #12

Earlier quoted context omitted.

> And while a single frame might be missing a label, I bet that at-speed most everything important gets labeled correctly enough to be better than a distracted human driver, or the average human driver for that matter. Would you bet a family member? That a distracted driver is a hazard does not mean other drivers are safe, or even saf er .

>Would you bet a family member? "think of the children!" Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'. Thousands of people are dying every day in cars. This is not something we need to wait for it to be perfect. It only need…

> Lives at stake don't change anything here. The question is whether self-driving cars, even with the errors, are safer for people than regular drivers on average. If so, then absolutely yes everyone should bet their lives and their families'.

Maybe logically that makes sense but from an ethical perspective I argue it's much more complicated than that (e.g. the trolley problem)

In the current system if a human is at fault, they take the blame for the accident. If we decide to move to self driving cars that we know are far from perfect but statistically better than humans, who do we blame when an accident inevitably happens? Do we blame the manufacturer even though their system is operating within the limits they've advertised?

Or do we just say well, it's better than it used to be and it's no one's fault? When the systems become significantly better than humans, I can see this perhaps being a reasonable argument, but if it's just slightly better, I'm not sure people will be convinced.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#39
post #11

It is a self-correcting problem: these pedestrains won’t be present in the next dataset.

People should learn not to go outside if they're not labelled.

What a funny future it'd be, if we have to wear something distinctive (giant QR codes?) so we don't get killed outside. As a side effect it would make tracking us much easier...

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#40
post #22

Earlier quoted context omitted.

It's scary not because this specific dataset was used to train Teslas that are on the road today. Rather, because it makes us aware of an entire class of errors that most of us probably hadn't thought about before. I guess you are absolutely certain that training data used in production cars will be free of these issues, but it's not clear why.

It does not make use aware of a new class of errors. Labeling issues is nothing new, but plenty of systems trained on them continue to work just fine. This is FUD.

Is there any statistical/mathematical tool to completely eradicate or greatly diminish the effects of bad labeling? Is there any reason - other than the combination of pure circumstance and gut feeling of the Data Scientist in charge of saying that it's good enough to deploy - that ~33% insanity in training doesn't become ~33% insanity in the system?
Post reply on HN