Earlier quoted context omitted.
You are reporting on a deliberately curated effort vs. what I understand is effectively voluntary data donation without incentives. It's not surprising to me that the later dataset ends up biased due to the differences in sourcing.
Your understanding of the datasets I helped create seems at odds with my experience actually creating the datasets. Do you have some insider experience or knowledge with dataset curation and creation for voice assistants that contradicts my own. The guideline is that the newer your model, the more likely it is to have diverse voice recognition datasets since it solves the earlier problems caused by non representative…
The effort being discussed is a volunteer effort among a community of tech enthusiasts, who are disproportionately privacy-oriented vs the average person. This will undoubtable skew towards middle-aged male audiences, and will be extra-selective against children. It's a best-effort collection, they're probably not turning anyone away, and it's only what they can get, they're (AFAIK) not paying anyone to collect underrepresented demographics.