Live data from Hacker News

A popular self-driving car dataset is missing labels for hundreds of pedestrians

blog.roboflow.ai

151–160 of 202 posts

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#151
post #102

Earlier quoted context omitted.

> No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space Are you sure about that? What about "non-serious" companies? Various fly-by-night self-driving startups? I mean, this sounds like the machine learning equivalent of "no serious business is pulling random bits of code from StackOve…

Are you sure about that? What about "non-serious" companies? Various fly-by-night self-driving startups? Fly-by-night self-driving startups will at some point run into issues with the NHTSA. George Hotz self-driving car project was shut down as soon as he announced that he'd start selling some prototype. States require permits just to be allowed to test self-driving prototypes on the open road. I don't know what the…

Given HN's general suspicion of government regulation, especially its quality and accuracy, I'm surprised to see people argue here that it's sufficient to counter the "move fast and break things" mentality. Especially given that regulatory arbitrage (moving from CA to AZ) was involved in Uber's vehicular manslaughter.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#153

Earlier quoted context omitted.

> People think that calling something "Artificial Intelligence" implies that it is artificial, yes, but also, critically that it is intelligent. I mean, the exact same argument could be made about people. Just because you belong to Homo Sapiens doesn't imply you're an intelligent being at all times - e.g. drink a bottle of vodka and the sapiens part is gone.

The thing with people is that we all share the same brain architecture - we mostly think alike. As a species, we have a hundred thousand plus years of experience dealing with each other; as a civilization, a couple thousand. We've explored most corner cases, designed our infrastructure around those, and built systems protecting us from outliers. For instance, people deemed too unpredictable will not be given driving…

I agree entirely with your point, and wanted to follow a small tangent:

> they lack self-preservation instinct

I suspect that a self-preservation instinct, even in relatively dumb animals, is enormously complex. It might seem simple to us because it's so evolved, so hardwired. If I had to bet, I'd bet that a mouse's self-preservation instinct is more complex than the first-to-market Level 5 self-driving car will be.

And that's good, because I definitely don't want automated vehicles to have more than a hint of self-preservation wiring.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#154

Earlier quoted context omitted.

That's probably the least charitable interpretation possible. No idea who the OP is, in any case. It's scary that the only way we know how to build something that detects pedestrians in an image with any kind of reliability is to use training data. You: All training data has labeling issues. Me: Training data is the only way we know how to build some aspects of systems. Other people here: Some of these systems are sa…

> No idea who the OP is, in any case. The person that originally posted this thread that works for the company that's selling this FUD. At the bottom of the blog post: "Roboflow accelerates your computer vision workflow through automated annotation quality assurance" This is a non-issue exasperated by a for-profit corporation that's creating an "issue" they conveniently have a service to help you fix.

I think you missed the key part of the post you're replying to:

> tell us why you feel so reassured that this is not a problem

Because to me you're coming across as the old-school "you don't need QA if your code is perfect" type.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#155

So, do they claim to be human managed datasets where people drew all the object bounds? I wonder if they used their own classification AI on sample sets and just called it a day. AI blind leading the AI blind?

They claim it was ("The dataset was annotated entirely by humans using Autti" via https://github.com/udacity/self-driving-car/tree/master/anno...).

But after looking at the data I'm almost certain there was some "tool assist" going on. There were dozens of frames in a row with phantom bounding boxes in the exact same location which made it look in some way automated (maybe user error combined with a "copy bounding boxes from the previous frame" feature?)

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#156

Earlier quoted context omitted.

> This is really scary. No, it's not even remotely "really scary". No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space or is represented of the impact a few mislabelled samples have on the ability of these systems & algorithms to generalize.

> No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space Are you sure about that? What about "non-serious" companies? Various fly-by-night self-driving startups? I mean, this sounds like the machine learning equivalent of "no serious business is pulling random bits of code from StackOve…

> Various fly-by-night self-driving startups?

What is the alternative other than holding these companies responsible for the deaths they might cause ? The only solution I see is mandating very high insurance requirement than a human driver. For example 25M per death they might cause.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#157

Earlier quoted context omitted.

Deleting half a sentence and saying it doesn't make sense? The way "average" works is that you average over something - a population or set. It is very important to be clear about what that something is and whether it's appropriate. Why do you believe that Autopilot outperforms humans in comparable conditions? If this is based on Tesla marketing, I'm extremely prejudiced against them, and assume out of hand that they…

The average of something cannot be more than the average of... itself. Thus, "Humans are much safer than people on average" is nonsensical. > Why do you believe that Autopilot outperforms humans in comparable conditions? Because they have the data that proves it? > I'm extremely prejudiced against them And I've chosen to take them at face value with a grain of salt, and to believe that for the data they've collected…

"Thus, "Humans are much safer than people on average" is nonsensical."

Does it make more sense as "Humans, when driving in conditions suitable for Autopilot, are much safer than people on average"?

"I've chosen to take them at face value with a grain of salt, and to believe that for the data they've collected from the hundreds of thousands of Tesla's with millions of hours of data using Autopilot, it's fair to say they have a large enough sample to draw conclusions about the safety of their cars vs. any incident rates from pretty much any other distribution"

You seem to be saying that if you have a lot of data it doesn't matter what you compare it to. That seems wrong to me. Also, I don't have this data, and you are not bothering to help me find it.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#158
post #2

This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…

> This is really scary. No, it's not even remotely "really scary". No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space or is represented of the impact a few mislabelled samples have on the ability of these systems & algorithms to generalize.

[deleted]

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#159
post #2

This is really scary. I discovered this because we're working on converting and re-hosting popular datasets in many popular formats for easy use across models... I first noticed that there were a bunch of completely unlabeled images. Upon digging in, I was appalled that fully 1/3 of the images contained errors or omissions! Some are small (eg a part of a car on the edge of the frame or a ways in the distance not bein…

> This is really scary. No, it's not even remotely "really scary". No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space or is represented of the impact a few mislabelled samples have on the ability of these systems & algorithms to generalize.

The scary part isn't necessarily this dataset, but that unlabeled data causes a silent decrease in model performance -- which can be esp important for underrepresented classes.

Re: A popular self-driving car dataset is missing labels for hundreds of pedestrians

#160

Earlier quoted context omitted.

> No one is putting an actual self-driving car on the market using this specific data set. Disingenuous to pretend this is any indication of the data using by serious companies in the space Are you sure about that? What about "non-serious" companies? Various fly-by-night self-driving startups? I mean, this sounds like the machine learning equivalent of "no serious business is pulling random bits of code from StackOve…

> Various fly-by-night self-driving startups? What is the alternative other than holding these companies responsible for the deaths they might cause ? The only solution I see is mandating very high insurance requirement than a human driver. For example 25M per death they might cause.

Perhaps reasonable regulations involving extensive pre-street testing and solid public documentation of ML design decisions and training data sources before allowing use in uncontrolled environments? Also, criminal liability during initial deployment?

Of course, real responsibility is probably unfashionable...

(edit: I'm personally in high-stakes time-sensitive diagnostic ML development, so...)

Post reply on HN