Author misses the way models work entirely, the larger the entity, the more statistics and averages kick in, and as a result, better model can be built.
Goldman Sachs model to predict World Cup game results didn’t come close
111–120 of 136 posts
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#112Earlier quoted context omitted.
So this is something that people don't seem to grok quite well, and it really depends on the type of statistical analysis used. Say you make the assumption that the quantity being estimated is truly fixed: that there's some true value for the force of gravity or some true value for the number of people that vote for X or Y. The second assumption that comes along is that the stochasticity observed comes from your pers…
If you expect to get it right (in this particular prediction, Clinton to win) with 95% probability, what does it mean to say that this 95% is with low confidence or with high confidence?
Basically, confidence refers to a hypothetical scenario in which a the data gathering process were to be repeated and the same analysis done, X% of the confidence intervals (essentially, the +/- bounds around your estimate) will contain the true value for what you are trying to estimate.
So in this hypothetical scenario, we say we have the power to go back in time and recollect the polling data in 2016 and run the same analysis used to arrive at that 95% number. And let’s say we use this power over and over again, a very large number of times. Then 95% of the error bounds we construct should contain the true value of the probability Hillary wins, whatever that is.
The thing is that those error bounds can be huge. You can have 95% confidence that the probability that Hillary wins is between 3% and 98%, for example. You can also have 10% confidence that the probability of a Hillary win is between 94% and 96%. Without the confidence intervals, a “confidence level” doesn’t say much. It’s also predicated on the assumption you haven’t screwed up your data collection process or analysis methodology. And if you are predicting something will occur with a probability of 95%, and it doesn’t, that doesn’t automatically mean you are wrong, but the likelihood of you having screwed something up is definitely higher.
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#113>Soccer, with the many factors that affect game outcomes — players’ injuries and intra-team conflicts, the refereeing, the weather, coaches’ errors and moments of inspiration — remains only a tightly-regulated game involving a few dozen people . The behavior and performance of big corporations, entire industries and nations is arguably even more difficult to model based on data about the past. Author misses the way m…
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#114Earlier quoted context omitted.
For example: Croatia got to the final via two penalty shootouts, which is very much like coin flipping.
Can't you put probabilities on that though? What you said is basically "impossible to predict, because one or the other might score more goals, so it's like coin flipping."
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#115People conflate statistics with actual results more often than not and I think those reporting on such stories and maybe even the original authors might fall for this. It was not wrong to say Hillary had a 95% chance of winning the presidential election, but the confidence was low and that value still allowed for the opposite result to happen . Also football has a lot of variance concerning team capability and end re…
> had a 95% chance [...] but the confidence was low So she had 95% chance of winning with 50% probability or what?
For example, if you are estimating the height of a male in the US, you would collect data on US males and get the average. But unless you surveyed every male in the US, there is some error associated with your estimate. So you would either construct error bounds (a frequentist approach) or a probability distribution (a Bayesian approach) around the mean height. So your results may dictate that the mean height of the American male is 5’11, plus or minus 2 inches. Those two inches represent uncertainty around your data collection. That’s the exact same thing that is done here, but with a percentage instead of a height. Outlets may predict Hillary winning at 95%, but the reality is their methodology should provide a plus-minus value around that. The problem is that few of them actually report that.
But it gets more confusing. That error bound is only around the mean. Pick a random guy out and not only will he likely not be 5’11, there is a decent chance he will be outside of that range of 5’9 - 6’1. You will get 5’7 guys and 6’4 guys pretty commonly. In the case of the election, it may actually be true that Hillary had a chance between say, 93% and 97% of winning. But even if that is the case, she will still lose between 3-7% of the time. But since we only have one reality to observe, we can’t know if she lost simply because we saw that 3-7% realized, or because they people coming up with that number screwed up. That’s why groups like 538 deserve more leeway. When they say that Donald Truml has a 30% chance of winning, and he does. That’s not that crazy. And therefore there is much less reason to assume they screwed something up than the people who predicted a 5% chance of Trump winning. It’s possible those models were right, but much less so.
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#116The "Ludic Fallacy" strikes again [0]. > The ludic fallacy, identified by Nassim Nicholas Taleb in his 2007 book The Black Swan, is "the misuse of games to model real-life situations." ... > The alleged fallacy is a central argument in the book and a rebuttal of the predictive mathematical models used to predict the future – as well as an attack on the idea of applying naïve and simplified statistical models in compl…
Damn, do I disliked Nassim Taleb. I don't think I've ever heard him say anything deep. That wikipedia article is an excellent. In [0] you have the following: > The ludic fallacy, identified by Nassim Nicholas Taleb in his 2007 book The Black Swan, is "the misuse of games to model real-life situations." And he gives an example of this: > One example given in the book is the following thought experiment. Two people are…
I’m not sure why there are so many people who take him seriously.
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#117Earlier quoted context omitted.
These kinda show what makes predicting football particularly difficult. I like the ideas, and I think we (or more likely some ML algorithm) can come up with the set of conditions that showed why France prevailed against the specific opposition at this specific World Cup ... but I suspect that the conditions would be pretty unique and invalid for Euro 2020, WC 2022 etc. As you identified, motivation could be pretty ha…
Thank you for the warm words, I guess the reason is my occupation plus the fact that I just spend my last few weeks watching many games with family and friends. > I'm not sure what you mean with the last one, but I think this could be a nice one - if you mean "times you lost possession in your own half" Almost, England lost the ball frequently (> 50+x% with a large x AFAI could see) due to the keeper sending out long…
Interestingly something like this is a tactic used in Rugby (https://www.youtube.com/watch?v=cbti6mLvSJs). I used to play a lot of football when I was younger and at our level (waaaay down the scottish league pyramid) against tired, hungover or generally weak opposition, keeping them under pressure by dominating the territorial game but sacrificing possession was criminally underrated. Usually if you could keep hammering them for 60 minutes and had the legs to step up a gear in the last 30 or so you could grab a valuable goal or two :-)
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#118Isn’t this a little like flipping a coin four times - getting heads four times in a row, and looking at your friend and saying “but you told me the odds were 50/50 each flip?!”
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#119Earlier quoted context omitted.
Their model also had France at 2nd most likely, Belgium at 5th, and England at 7th. 3 of their top 7 made the Semi-Finals, and they called the eventual winner as Second Most Likely, and more likely than Germany. They actually predicted the Brazil/Belgium game in the Quarter Finals, but got the winner wrong. Brazil had 27 shots and 9 on target with 59% posession. Belgium only had three shots on target, and made two of…
> Brazil had 27 shots and 9 on target with 59% posession. Belgium only had three shots on target, and made two of them to win. A modern model would accommodate for the fact that those numbers alone mean nothing, because they don't. Those are the numbers broadcasters reluctantly put on a screen for entertainment value, but they don't have real analytical power because they have no comparative metric. How up or down we…
You're not going to find a statistical approach that will account for the subtleties that led to this outcome. The problem with soccer stats in general is that everything hinges on low-frequency events based on subtle differences of timing and space.
Basketball by comparison is much more stat-rich, and there are a lot of cool advanced analytics, but even still they are full of gaps that are obvious to any expert watching the game. Afterwards maybe you can find the statistical signature of something you saw, but then you risk overfitting again, just the same as soccer.
Re: Goldman Sachs model to predict World Cup game results didn’t come close
#120Earlier quoted context omitted.
> The model said, that there is a lot of uncertainty, and as it happens, it was entirely correct. A World Cup chance of 18.5 percent means, that 4 out of 5 times the team will not win, and that that is the highest chance does not say much about the model. But do you need a sophisticated model and lots of so-called "AI" to arrive at the conclusion that there's a lot of uncertainty?? The point of the model is to reduce…
The point of the model is absolutely not to reduce uncertainty, it is to quantify it, which are two very different things. No model reduces uncertainty in a probabilistic sense. And no, you don’t need statistics or machine learning to say “there is a lot of uncertainty”, but you do in order to quantify that uncertainty.