Live data from Hacker News

When the U.S. air force discovered the flaw of averages

thestar.com

81–90 of 115 posts

Re: When the U.S. air force discovered the flaw of averages

#81
post #59

While hindsight is 20:20, parts of this should have been more obvious. > Daniels generously defined as someone whose measurements were within the middle 30 per cent of the range of values > Daniels discovered that if you picked out just three of the ten dimensions of size ... less than 3.5 per cent of pilots would be average sized on all three dimensions. 30% raised to the third power is 2.7%. Basic probability. I gu…

That is over-simplification. First, the samples are not randomly selected. The selection process tended to lean toward means, based on the article. Second, the variables are not independent. They are strongly correlated. The selection process indicated that the probability of having a pilot with long legs but short arms was close to zero.

> They are strongly correlated

Uh, no. What the study shows is that people assumed these measurements would be strongly correlated, but they were not.

When only 3.5% of people fall into the middle 30% on 3 variables, that is strong evidence that they are mostly independent.

Re: When the U.S. air force discovered the flaw of averages

#82
post #11

Earlier quoted context omitted.

Yeah, in this case if the 10 traits are independent and the chance of falling in the "average range" is 1/3 for any one trait then the probability of any one soldier falling the the average range for 10 traits is (1/3)^10 = 1/59049.

That is true, but this can be just as much an issue in 1 dimension. Consider that the average person has roughly 1 testicle. I think it's about understanding distributions, and joint distributions are a big part of that.

[deleted]

Re: When the U.S. air force discovered the flaw of averages

#83

Earlier quoted context omitted.

> A cockpit is not a suit. The opposite is true. To fly an airplane with precision, you need to "strap the airplane on." That means sitting in the position with the best visibility and best reach inside the cockpit, even a Cessna 172. (In a 172, you should sit high enough to be able to see each of the rivets on top of the cowling.)

Yes, but the seat does adjust. It isn't the case that the cockpit must be tailored to every person. Each Cessna, 747 or f-16 seat is practically interchangeable amongst the type. These aren't WWI biplanes with nothing more than a cushion. Even then, the cushion height, the prime dictator of eye level, is adjustable.

> Yes, but the seat does adjust.

It does now. It didn't then.

Did you read the article? That's the entire point of it!

Re: When the U.S. air force discovered the flaw of averages

#84

Wow! This has implications for human genetics (a study that originated partly in considering human body size measurements like those described in the article) that I think most popular writers on human genetics have not taken into account. There was an article posted to Hacker News earlier this week that suggested, surely falsely in my opinion, that human genomes can be optimized for intelligence to such a degree tha…

You might like this article: http://blog.watsi.org/netflix-and-global-health/

Great article.

Still the Netflix recommendations are awful and contradict my taste whatever I tried to help it understand the movies/series I liked and disliked. As result I subscribe one month, this month, to see House of Cards and then unsubscribe till next season. Maybe this average algorithm works for illnesses and mortality, IMHO it is not biased enough to my non-US taste/culture to be of any use for Netflix recommendations.

Re: When the U.S. air force discovered the flaw of averages

#85

Earlier quoted context omitted.

I don't think that's quite correct. Recommendation engines normally work by taking the things you favor, then looking at who else favors those things, and deducting that you're all likely to have shared interests.

Yeah I'd assume the simplest recommendation engines work on correlation. A recommendation engine working on average would be complete garbage, you'd be better off picking recommendation at random.

Quote emcq's response above:

> If you're thinking about the cold start problem when you dont have any information about a user, yes it's possible that your overall statistics is a combination of many subpopulations that doesn't really fit anyone very accurately but there are ways around this as well.

I'd say a better approach (but may have user experience penalty) would don't even decide for the user, ASK. For example, take Flipboard / Quora as an example, you are asked to choose some topics to follow at the beginning. Assume research/data show 80% of the users are software engineers and 90% of them always pick "technology" as a topic they want to follow, would you rather show technology as one of the top five in the list of topics to choose? There's actually a lot of experiments you can do from a simple selection/survey process. I personally can't stand at going through pages to find something relevant, but I am also surprised to find things I never thought would be interesting to follow if I weren't present the options at all / or earlier.

Re: When the U.S. air force discovered the flaw of averages

#87

Wow! This has implications for human genetics (a study that originated partly in considering human body size measurements like those described in the article) that I think most popular writers on human genetics have not taken into account. There was an article posted to Hacker News earlier this week that suggested, surely falsely in my opinion, that human genomes can be optimized for intelligence to such a degree tha…

Similarly, imputation of "race" by genome testing depends crucially on assuming that the average genome is informative, This is a nonsensical mischaracterization of both modern genetics and the theory of high dimensional vector spaces. The phenomenon the article is discussing is the fact that the mass near the center of a normal distribution approaches zero as the number of dimensions goes up, and that most of the ma…

I can talk about some of this, if only because of some of the ethical dilemmas facing me and genome testing as of recent.

Racial Groups as described by a geneticist vs say, a government doing polling on its citizens are different animals.

Goverments == mostly sociological. An example would be Hispanic in the US, where growing up data collection would ask if someone was Hispanic, and now it asks if you are white vs non-white Hispanic. Meanwhile in say, Meanwhile, if someone crossed the border to Mexico, there is no such thing as Hispanic - you can be White, Mestizo (Indigenous-European hybrid), Indigenous, and Other. It is totally possible to live on the border of the US-Mexico, have reasons to commute across the US-Mexican border, have citizenship to both countries, and have totally different answers on your census depending on what country asked you about your race/ethnicity, because as can be clearly seen, the way that question is asked is different in mexico and the US to begin with.

Let's talk about being Mezito as an idea. From a geneticist's point of view, it doesn't exist. (or at least, not yet, and it is unliekly to any time soon) Which is how 23 and me and buzzfeed manages to get this Gif off of one of Buzzfeed's employees, who very clearly feels he is half mexican (aka mezito) https://img.buzzfeed.com/buzzfeed-static/static/2016-01/21/1... In other words, there are no clusters that define "mexican"/mezito

In order to have such clusters, you need to have long histories in one area, histories of inbreeding, and other major causes to make mutations pop selectively. Even with those mutations popping selectively, you also need those mutations/genes to be very trait specific in most cases/ultra selective. Otherwise, you are looking at junk.

So, one of the reasons ashkenazim are testable as ashkenazim is because they have a very long history of inbreeding. So do the japanese compared to other asian groups, same with the Finnish, especially if you are Saami.

http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2575505/

Still once these ibreeding/long history facts start to subside, you end up with whatever is the strongest trait set in a pure mendelian fashion. It is why there are plenty of ashkenazim who definitely decendants of the founder populations of ashkenazim who don't carry other sets of Ashkenazi marker genes, whereas there are ashkenazim where you can't tell who do carry these other genes. The other random sets come from other places (sex..mutations..) and are bred out with enough time. (Otherwise there would be many many more questions about how I exist, genetically speaking, since I should not be able to metabolize alcohol in about 1.5 hours...)

However, most traits are not one gene == one trait. It's how the music plays together, especially in concert with epigentic factors. For many things involving "intelligence" (or actually lots of things for humans), we're looking at a non-mendellian inheritance pattern, including Genomic imprinting issues ( https://en.wikipedia.org/wiki/Genomic_imprinting ). So while some gene related to "intelligence"/"thinness"/"insert something here" might be more observable among certain groups/populations, it doesn't mean it doesn't exist among other populations and some other factor, genetic, epigentic, or otherwise, is suppressing your view.

This is why if/when you get involved with genetic research on humans, the ideal is to get the entire family (as much as possible) involved in tests - even if the original candidate comes from a human genetic isolate polpulation. There are otherwise broader concerns about power and sample size in the study, because you might be looking at something which is actually something else.

http://www.nature.com/gim/journal/v4/n2/full/gim200210a.html

(but, hey, what do I know...I had to ask geneticists and people doing research in this area when I found out my genome is worth something accidentally...)

Re: When the U.S. air force discovered the flaw of averages

#88
post #59

While hindsight is 20:20, parts of this should have been more obvious. > Daniels generously defined as someone whose measurements were within the middle 30 per cent of the range of values > Daniels discovered that if you picked out just three of the ten dimensions of size ... less than 3.5 per cent of pilots would be average sized on all three dimensions. 30% raised to the third power is 2.7%. Basic probability. I gu…

> The Aero Medical Laboratory hired Daniels because he had majored in physical anthropology, a field that specialized in the anatomy of humans, as an undergraduate at Harvard. During the first half of the 20th century, this field focused heavily on trying to classify the personalities of groups of people according to their average body shapes — a practice known as “typing.” For example, many physical anthropologists…

Lombroso et al.

Worth noting that this phenomenon is typically described as “medicalization.” Othered groups medicalized include women, African Americans, homosexuals, and many other groups. The medicalization of Jewish ancestry cropped up and was amplified during a not-too-distant past.

Re: When the U.S. air force discovered the flaw of averages

#89
post #59

While hindsight is 20:20, parts of this should have been more obvious. > Daniels generously defined as someone whose measurements were within the middle 30 per cent of the range of values > Daniels discovered that if you picked out just three of the ten dimensions of size ... less than 3.5 per cent of pilots would be average sized on all three dimensions. 30% raised to the third power is 2.7%. Basic probability. I gu…

> The Aero Medical Laboratory hired Daniels because he had majored in physical anthropology, a field that specialized in the anatomy of humans, as an undergraduate at Harvard. During the first half of the 20th century, this field focused heavily on trying to classify the personalities of groups of people according to their average body shapes — a practice known as “typing.” For example, many physical anthropologists…

[deleted]

Re: When the U.S. air force discovered the flaw of averages

#90
post #39

Earlier quoted context omitted.

Yes, this can be made precise: For example, consider a uniform distribution on a d-dimensional cube with side length 1. Only a (1/2)^d fraction of the mass is within 0.25 distance of the average in every coordinate. In high dimensions, almost all the mass is "near the boundary" in at least one coordinate.

Or as one of my professor would say (after showing the calculation for n-dimensional spheres): "and that's why infinite dimensional oranges are all skin!"

This must be a common saying. I've heard it from every lecturer I've ever had for machine learning and sampling techniques.
Post reply on HN