> First, to recap: Google injected special instructions into Gemini so that when it was asked to draw pictures, it would draw people with “diverse” (non-white) racial backgrounds. I just love how this "solution" to their problem of their training set being deeply flawed is such a hilariously bad kludge hack of the sort a jr programmer would make and it's amazing that they went with it.
The world is deeply flawed. If you do an unbiased search [0] of the real pictures of people in such-and-such profession, you get a vague approximation of the people actually in that profession. This does not approximate what Google or anyone else wants the demographic to be — it’s the demographic of people matching the query, times the expected number of times their photographs appeared in the training set, times whatever weight the model gives to those photographs. The naturally picks up whatever may be wrong with society that resulted in this distribution.
If people of X race/gender/whatever are underrepresented in profession Y, and you search for photos (or synthesize them by generative AI), do you want a sample drawn from the actual distribution of photos or from a distribution with the underrepresentation corrected? And what does correcting it mean?
[0] whatever that means — maybe just giving uniform weight to all the distinct pictures of decent quality that are on the Internet. The actual selection criteria won’t change the result too much.