Please be more careful when interpreting the Stack Overflow Developer Survey
meta.stackoverflow.com
Please be more careful when interpreting the Stack Overflow Developer Survey
1–10 of 22 posts
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#2Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#3Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#4So, honest question:
If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful?
Isn’t this a self defeating argue?
You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent... but how do we go away and prove that some other survey we’re comparing it to is from a random sample?
ie. Isnt this just a convenient excuse to deny that a survey is meaningful?
Statistically, how do you mathemtaically quantify the effect of selection bias?
...because, it seems to me, unless you can actually do that, you’re just doing some arm chairmhand waving because you don’t like the results youre seeing.
This has come up several times (eg. js survey about react vs angular), and no one has ever given me a meaningful and mathematical response.
Its always just.. “it must be sample bias”, regardless of the 90000 people they surveyed.
I don’t accept you can survey 90000 developers and cannot offer any generalisation from those results without quanatitively proving there is an overwhelming sample bias, and specifically quantifying the degree of that bias.
Am I missing something here? Everyone seems thoughorly convinced that this is perfectly normal.
(I’m not proud, I’ll take your down votes, but please answer and explain what I’m missing)
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#5> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#6> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
Surely you have this backwards? If you want to argue that a survey offers any generalisation, then surely the onus is on you to prove you've accounted for sample bias (amongst others)?
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#7> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
This survey obviously does not tap directly into the brains of every developer on the planet and extract their unbiased answers to the questions. But it’s still a useful model for seeing trends in the software industry.
Further, from my personal perspective, I’m pretty ok with the self-selected sampling bias inherent in the survey. The kinds of developers who see the value of Stackoverflow and are willing to participate voluntarily in the survey are the kinds of developers whose opinions I generally care about :). That’s my own bias, which I acknowledge exists, and it doesn’t particularly bother me.
Edit: further, none of the results jump out at me as particularly surprising. If there were some extraordinary results here, I’d want someone to do a more rigorous follow-up to dig into that, but there isn’t so...
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#8> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
https://www.math.upenn.edu/~deturck/m170/wk4/lecture/case1.h...
The way to deal with this is to try to construct a representative sample. Here is Gallup's method in 1936:
> But George Gallup knew that huge samples did not guarantee accuracy. The method he relied on was called quota sampling, a technique also used at the time by polling pioneers Archibald Crossley and Elmo Roper. The idea was to canvass groups of people who were representative of the electorate. Gallup sent out hundreds of interviewers across the country, each of whom was given quotas for different types of respondents; so many middle-class urban women, so many lower-class rural men, and so on. Gallup's team conducted some 3,000 interviews, but nowhere near the 10 million polled that year by the Digest.
Stack Overflow did not attempt to construct a representative sample of developers. Therefore they cannot claim that we can learn from their sample about the population.
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#9> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
Nope, not at all.
It is true that no sample will ever be perfectly representative of a larger population. However, some samples can clearly be more representative than others and the easiest way to tell is to look at sampling methodology. Sample size has absolutely no effect on removing bias.
Here's some info so you can learn more about sampling methodologies: https://blog.socialcops.com/academy/resources/6-sampling-tec...
Now, doing actually representative samples is HARD in many situations, so is knowing how representative your sample is. This is why we can't predict things like who is going to win elections.
Re: Please be more careful when interpreting the Stack Overflow Developer Survey
#10> you cannot generalize from a non-random sample So, honest question: If any survey of any size can be ignored on the basis that the sample is not random, then how is any survey meaningful? Isn’t this a self defeating argue? You can’t prove the sample is random, all you can do is show differences between samples and suggest its not consistent ... but how do we go away and prove that some other survey we’re comparing…
> I don’t accept you can survey 90000 developers and cannot offer any generalisation from those results without quanatitively proving there is an overwhelming sample bias, and specifically quantifying the degree of that bias. Surely you have this backwards? If you want to argue that a survey offers any generalisation, then surely the onus is on you to prove you've accounted for sample bias (amongst others)?
If you want to argue with it, surely the onus is on you to do it concretely?
> Because of your methodology, we must assume a biased sample.
^ I find this quote problematic.
Why must we assume that? If you want to distribution comparisons and point out there survey results are skewed by X compared to some other survey Y... ok.
...but that’s not whats happening right? Its just a flat out arbitrary assumption.
I don’t like arbitrary assumptions when I’m doing maths.
Its easy to say something is wrong, but if you can’t quanitfy how its wrong, I’m struggling to see why I should accept the assumption being raised here.
The js survey was very similar; it was arbitrarily asserted it went to more react developers... but no one actually proved that. They just... assumed it.