So Crawford doesn't actually ever assert that MTurk is used directly for Alexa, only that it's one common form of commoditized data labor, of which the internal teams of annotators is actually another good example.

The language thing is even more interesting. If you read Crawford's research, her whole thing is about how machine learning is full of bias, for example about how AmazonFresh availability maps resemble 1930s segregation maps. Yes, it gets kind of philosophical on the discussion, and I'm nowhere near qualified enough to get into a discussion about this, but the thing that I'm curious is, how much does the data reflect the biases of those language engineers?