Live data from Hacker News

Please be more careful when interpreting the Stack Overflow Developer Survey

meta.stackoverflow.com

1–10 of 22 posts

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#4
> you cannot generalize from a non-random sample

So, honest question:

If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful?

Isn’t this a self defeating argue?

You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent... but how do we go away and prove that some other survey we’re comparing it to is from a random sample?

ie. Isnt this just a convenient excuse to deny that a survey is meaningful?

Statistically, how do you mathemtaically quantify the effect of selection bias?

...because, it seems to me, unless you can actually do that, you’re just doing some arm chairmhand waving because you don’t like the results youre seeing.

This has come up several times (eg. js survey about react vs angular), and no one has ever given me a meaningful and mathematical response.

Its always just.. “it must be sample bias”, regardless of the 90000 people they surveyed.

I don’t accept you can survey 90000 developers and cannot offer any generalisation from those results without quanatitively proving there is an overwhelming sample bias, and specifically quantifying the degree of that bias.

Am I missing something here? Everyone seems thoughorly convinced that this is perfectly normal.

(I’m not proud, I’ll take your down votes, but please answer and explain what I’m missing)

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#5

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

Anybody can choose to ignore anything. But, if the selection process is demonstrably random (no accidental biases) and we can assume that the SO people are trustworthy (no intentional biases), we can then start making generalizations about the entire population. Everything ultimately comes down to trust, unless you run your own study. But then, do you trust yourself to conduct it properly?

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#6

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

> I don’t accept you can survey 90000 developers and cannot offer any generalisation from those results without quanatitively proving there is an overwhelming sample bias, and specifically quantifying the degree of that bias.

Surely you have this backwards? If you want to argue that a survey offers any generalisation, then surely the onus is on you to prove you've accounted for sample bias (amongst others)?

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#7

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

I’m with you. How does the saying go? “All models are flawed, some models are useful.”

This survey obviously does not tap directly into the brains of every developer on the planet and extract their unbiased answers to the questions. But it’s still a useful model for seeing trends in the software industry.

Further, from my personal perspective, I’m pretty ok with the self-selected sampling bias inherent in the survey. The kinds of developers who see the value of Stackoverflow and are willing to participate voluntarily in the survey are the kinds of developers whose opinions I generally care about :). That’s my own bias, which I acknowledge exists, and it doesn’t particularly bother me.

Edit: further, none of the results jump out at me as particularly surprising. If there were some extraordinary results here, I’d want someone to do a more rigorous follow-up to dig into that, but there isn’t so...

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#8

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

It IS sample bias. Read the linked article, which describes this as the worst kind of selection bias, when the sample is made of volunteers.

https://www.math.upenn.edu/~deturck/m170/wk4/lecture/case1.h...

The way to deal with this is to try to construct a representative sample. Here is Gallup's method in 1936:

> But George Gallup knew that huge samples did not guarantee accuracy. The method he relied on was called quota sampling, a technique also used at the time by polling pioneers Archibald Crossley and Elmo Roper. The idea was to canvass groups of people who were representative of the electorate. Gallup sent out hundreds of interviewers across the country, each of whom was given quotas for different types of respondents; so many middle-class urban women, so many lower-class rural men, and so on. Gallup's team conducted some 3,000 interviews, but nowhere near the 10 million polled that year by the Digest.

Stack Overflow did not attempt to construct a representative sample of developers. Therefore they cannot claim that we can learn from their sample about the population.

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#9

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

> ie. Isnt this just a convenient excuse to deny that a survey is meaningful?

Nope, not at all.

It is true that no sample will ever be perfectly representative of a larger population. However, some samples can clearly be more representative than others and the easiest way to tell is to look at sampling methodology. Sample size has absolutely no effect on removing bias.

Here's some info so you can learn more about sampling methodologies: https://blog.socialcops.com/academy/resources/6-sampling-tec...

Now, doing actually representative samples is HARD in many situations, so is knowing how representative your sample is. This is why we can't predict things like who is going to win elections.

Re: Please be more careful when interpreting the Stack Overflow Developer Survey

#10
post #6

> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…

> I don’t accept you can survey 90000 developers and cannot offer any generalisation from those results without quanatitively proving there is an overwhelming sample bias, and specifically quantifying the degree of that bias. Surely you have this backwards? If you want to argue that a survey offers any generalisation, then surely the onus is on you to prove you've accounted for sample bias (amongst others)?

That seems fair; but they have a whole methodology section.

If you want to argue with it, surely the onus is on you to do it concretely?

> Because of your methodology, we must assume a biased sample.

^ I find this quote problematic.

Why must we assume that? If you want to distribution comparisons and point out there survey results are skewed by X compared to some other survey Y... ok.

...but that’s not whats happening right? Its just a flat out arbitrary assumption.

I don’t like arbitrary assumptions when I’m doing maths.

Its easy to say something is wrong, but if you can’t quanitfy how its wrong, I’m struggling to see why I should accept the assumption being raised here.

The js survey was very similar; it was arbitrarily asserted it went to more react developers... but no one actually proved that. They just... assumed it.

Post reply on HN