Earlier quoted context omitted.
According to her, Yann LeCun is a racist just because he said that a specific ML model was biased due to a bias in the data and not because ML learning as a research field was racist. Such an inspiration!
She didn't call him a racist. She pointed out that there's more to addressing bias than just acknowledging that it's there. If you are going to say "Well, garbage-in, garbage-out!", why do you keep putting the racist garbage in? Why do you, knowing of the problem, keep using biased sources of data, knowing of the potential harms? That's the deeper question ML research/industry has to grapple with. When these models a…
Probably because it was a research AI not a production AI. Having a very diverse dataset at that stage doesn’t help with your research, so it is fine to use whatever is easily available.
At that stage, you are trying to show that your approach can work in some cases. Once you've got that, it is time to expand the research with a wider range of inputs to find out what the limits of your approach are.
For example, if I were trying to make a US English speech to text transcription system, I might start with recordings of assorted NPR programs, because NRP often makes the recording available online along with they transcripts.
That would be great for determining if my basic approach has promise. Once I have determined that, so know that the whole endeavor is not just a waste of time, I could go looking for data that includes speech that has characteristics that would be missing from the NPR data, such as heavy regional accents.