Live data from Hacker News

Predicting Football Results with Statistical Modelling

dashee87.github.io

21–30 of 32 posts

Re: Predicting Football Results with Statistical Modelling

#21

Remember, no one wins at gambling by picking winners, you need to look for value and find where the bookmakers have miss-priced a team in a match. Then you have to deal with all the corrupt behaviours the sports internet bookies will deploy to limit their exposure to you, which is the other reason why you won’t win.

Professionals don't use retail bookies, they bet with Asian books who take action from sharps (Pinny, SBOBet, SingBet, IBC...although it varies, sometimes their lines are soft).

Re: Predicting Football Results with Statistical Modelling

#22
post #7

The same author does seem to go on and build a slightly more realistic Dixon-Coles model: https://dashee87.github.io/football/python/predicting-footba... But even that is very dated - one of the authors, Stuart Coles has worked at a gambling syndicate for many years. They at the very least incorporate expected goals into the model, but no doubt have all sorts of esoteric models by now.

Yep, Dixon-Coles is way out now. Generally, Poisson-based models don't work that well. They model some aspects of the game correctly but not others. And when it is bad, it is terrible (their paper was genius though, and changed the industry). You also wouldn't do the one-hot/regression stuff on teams, as there is so much player-level data...some kind of online, off/def model is useful though. You also have the issue…

Any pointers or papers for someone who's just getting into this?

Re: Predicting Football Results with Statistical Modelling

#23
post #7

The same author does seem to go on and build a slightly more realistic Dixon-Coles model: https://dashee87.github.io/football/python/predicting-footba... But even that is very dated - one of the authors, Stuart Coles has worked at a gambling syndicate for many years. They at the very least incorporate expected goals into the model, but no doubt have all sorts of esoteric models by now.

Yep, Dixon-Coles is way out now. Generally, Poisson-based models don't work that well. They model some aspects of the game correctly but not others. And when it is bad, it is terrible (their paper was genius though, and changed the industry). You also wouldn't do the one-hot/regression stuff on teams, as there is so much player-level data...some kind of online, off/def model is useful though. You also have the issue…

The thing is, the size of the wager and the payouts are just as important. I was never a sports gambler, but spent two years counting cards at blackjack in rural casinos as my job. This plays out with the Kelly Criterion: https://en.wikipedia.org/wiki/Kelly_criterion So you can have inefficient odds as the house, and still win. Did the models also provide optimal bet size given the probability of winning?

Re: Predicting Football Results with Statistical Modelling

#24

Earlier quoted context omitted.

Yep, Dixon-Coles is way out now. Generally, Poisson-based models don't work that well. They model some aspects of the game correctly but not others. And when it is bad, it is terrible (their paper was genius though, and changed the industry). You also wouldn't do the one-hot/regression stuff on teams, as there is so much player-level data...some kind of online, off/def model is useful though. You also have the issue…

Any pointers or papers for someone who's just getting into this?

There's probably some interesting work to be done more from a survival analysis angle. I know I haven't seen much work based on stuff like:

https://www.academia.edu/37585525/Flexible_Regression_Models...

I do generally find these top-down team-strength models quite primitive though. Teams are made up of players, plans, all sorts of under-studied interactions. I would also say that there are enough quant jobs in football now that you don't really need to gamble to make a living. :)

Re: Predicting Football Results with Statistical Modelling

#25

Earlier quoted context omitted.

Yep, Dixon-Coles is way out now. Generally, Poisson-based models don't work that well. They model some aspects of the game correctly but not others. And when it is bad, it is terrible (their paper was genius though, and changed the industry). You also wouldn't do the one-hot/regression stuff on teams, as there is so much player-level data...some kind of online, off/def model is useful though. You also have the issue…

The thing is, the size of the wager and the payouts are just as important. I was never a sports gambler, but spent two years counting cards at blackjack in rural casinos as my job. This plays out with the Kelly Criterion: https://en.wikipedia.org/wiki/Kelly_criterion So you can have inefficient odds as the house, and still win. Did the models also provide optimal bet size given the probability of winning?

Larger syndicates also have the problem that it's not always easy to wager as much money as they'd like on a small handful of leagues. So it becomes about how you can accumulate data and insight into a broader set of competitions.

Re: Predicting Football Results with Statistical Modelling

#26

Earlier quoted context omitted.

Yep, Dixon-Coles is way out now. Generally, Poisson-based models don't work that well. They model some aspects of the game correctly but not others. And when it is bad, it is terrible (their paper was genius though, and changed the industry). You also wouldn't do the one-hot/regression stuff on teams, as there is so much player-level data...some kind of online, off/def model is useful though. You also have the issue…

Any pointers or papers for someone who's just getting into this?

Dixon-Coles, Ntzoufras, McHale, there are a few papers on Bivariate models from German authors (Groll...so google "Groll bivaraite Poisson"). I would also understand ranking algorithms (there is a book called Who's #1?). Bear in mind though, most papers are fictional/p-hacked/just bad.

I would also think carefully about what you are doing and why. Most people fail because they try to bet on markets that are, for them, unbeatable (for example, football data is expensive). It is far easier to pick off obscure markets (I did not do this because I had a junior high school maths education and needed some guidance).

Re: Predicting Football Results with Statistical Modelling

#27
post #25

Earlier quoted context omitted.

The thing is, the size of the wager and the payouts are just as important. I was never a sports gambler, but spent two years counting cards at blackjack in rural casinos as my job. This plays out with the Kelly Criterion: https://en.wikipedia.org/wiki/Kelly_criterion So you can have inefficient odds as the house, and still win. Did the models also provide optimal bet size given the probability of winning?

Larger syndicates also have the problem that it's not always easy to wager as much money as they'd like on a small handful of leagues. So it becomes about how you can accumulate data and insight into a broader set of competitions.

A lot of the newer syndicates spread themselves quite thin afaik. They maybe don't have the contacts to get liquidity/early prices so they do lots of sports (Tennis and Cricket being two that appear to be growing...again, afaik).

I have heard that the largest syndicate has groups that only cover one football team. I have no idea whether this is true or how/where you get the liquidity to justify this.

Re: Predicting Football Results with Statistical Modelling

#28
post #25

Earlier quoted context omitted.

The thing is, the size of the wager and the payouts are just as important. I was never a sports gambler, but spent two years counting cards at blackjack in rural casinos as my job. This plays out with the Kelly Criterion: https://en.wikipedia.org/wiki/Kelly_criterion So you can have inefficient odds as the house, and still win. Did the models also provide optimal bet size given the probability of winning?

Larger syndicates also have the problem that it's not always easy to wager as much money as they'd like on a small handful of leagues. So it becomes about how you can accumulate data and insight into a broader set of competitions.

This is the reality of how the bettors move the lines (ie odds on offer), large bets or a lot of small bets will move the line. In effect the final line at kickoff (or whatever) is kind of a distilled, crowd sourced expectation for the result of the match. The key is to identify advantageous lines when the major books publish them and place bets before the "public" moves the line.

Re: Predicting Football Results with Statistical Modelling

#29
post #2

Perhaps include "(soccer)" in the title for those of us in the USA?

Not the rest of the world's problem you called a game that mainly involves throwing 'foot' ball for some incomprehensible reason. Are you sure about the name basketball? Don't want to call it sackball or something like that? Or maybe shuttlebasket? I always chortle at your "world" series too!

I was going to respond with the factoid about the World Series being named after a newspaper, but a Wikipedia check (and subsequent googling) revealed that it's probably false. Not news to you, but maybe of interest to others who, like me, had absorbed the popular misconception!

Re: Predicting Football Results with Statistical Modelling

#30
post #2

Perhaps include "(soccer)" in the title for those of us in the USA?

Not the rest of the world's problem you called a game that mainly involves throwing 'foot' ball for some incomprehensible reason. Are you sure about the name basketball? Don't want to call it sackball or something like that? Or maybe shuttlebasket? I always chortle at your "world" series too!

American Football is actually based on the delineation between sports played on horseback and on foot. So maybe it is an antiquated distinction but it makes sense. Maybe we can rename it infantryball.
Post reply on HN