This is an interesting article, but there's one part that I don't really understand: 'After scaling up the approach by increasing the dataset size, we de-anonymize 600 programmers with 52% accuracy.'. Isn't 52% close to a random guess? I don't get how this ties in with the rest of the paragraph either.
If the question is, which of these 20 programmers wrote this code", then a random guess has only a 5% chance (1/20) of being right! So 52% is more than 10X better than random.