Live data from Hacker News

Frequentists should more often consider using Bayesian methods

thestatsgeek.com

11–20 of 39 posts

Re: Frequentists should more often consider using Bayesian methods

#11
Maybe it's because my formal math training is not in probability and statistics, but it's so bizarre to me that in a technical situation people would let a philosophical position dictate their approach rather than best tools for the job.

Sometimes I'll solve a math problem analytically, and sometimes its easier to do it numerically. But it would be foolish for me to take a hardline stance on one vs the other. Rather I am a more well-rounded, and thus more capable technician because I know the benefits and weaknesses of each approach.

If you have small-to-medium sample sizes, then clearly Fisher style statistics will not work very well. On the other hand, even if you have a small sample size, if you don't have some general knowledge to guide your priors, you may very well end up with garbage in a Bayesian approach.

Re: Frequentists should more often consider using Bayesian methods

#12
Many people learn about the philosophy of the Bayesian estimation and fall in love with it, or at least that happened to me.

I thought that only in the Bayesian formulation of statistics the estimated parameters (mean, standard deviation, kurtosis, percentiles, whatever) remain uncertain after the estimation, and therefore they have a distribution; I didn't know that in the frequentist interpretation what you calculate are actually estimators, and they are random variables and therefore have uncertainty in them. Silly me.

One day, reading about estimators on stackexchange, I was led to a quote from the "Elements of Statistical Learning" ([1], p. 272): "we might think of the bootstrap distribution as a poor man's Bayesian posterior".

So, I put off learning all the details of the Bayesian estimation, and I started using the bootstrap to deal with my problems. I recommend everyone to give the bootstrap method a try before they go all in for Bayesian estimation.

Once you get familiar with the bootstrap, you might feel that actually Bayesian estimation might be overkill.

But you shouldn't stop there. With a little more contemplation, you might start doubting the whole Bayesian edifice. Let me tell you why.

Here's a quote from a book [2] "Bayesian Risk Management" that, unsurprisingly, given the title, extols the virtues of the Bayesian framework: "If the data are consistent with our prior estimates, the location of the parameters will be little changed and the variance of the posterior distribution will shrink. If the data are surprising given our prior estimates, the variance will increase and the location will migrate" (p. 11). This is the common view of the Bayesian estimation, and it's wrong. To give you an example, if you do Bayesian estimation for the mean of a random sample assumed to come from a normal distribution with a given standard deviation, then the variance of that mean keeps going down with each observation regardless of the value of the observation (check first entry in [3])

But that shouldn't shatter your hope for a better world. The Bayesian estimation is bound to produce the wrong result if you assume the wrong model, regardless of what prior you use. But now you discover that the Bayesian estimation doesn't absolve you of the responsibility of choosing a good model (which is another commonly held believe). But then, if you do need to carefully choose a model, why do you need Bayesian after all.

I am not sure. Let me kick Bayesian while it's down a bit more, before I start defending it.

If you ever get curious about Kalman filter estimation, a good book to use is Durbin and Koopman [4]. In the first 20 pages you will learn that in the simplest setting (local level model), the Kalman filter gives exactly the same results in the frequentist and bayesian interpretation. Food for thought.

Here's a quote from Efron&Hastie "Computer Age Statistical Inference": "Computer-age statistical inference at its most successful combines elements of the two philosophies, as for instance the empirical Bayes methods in Chapter 6, and the lasso in Chapter 16. There are two arrows in the statistician's philosophical quiver, and faces, say, with 1000 parameters and 1,000,000 data points, there's no need to go hunting armed with just one of them."

Now I did say I'll come to the defense of Bayesian. To my knowledge the Markov Chain Monte Carlo method was developed only in the Bayesian setting. I do believe a frequentist interpretation is entirely possible, but so far nobody offered it. But until someone needs to use MCMC, I don't really see a need to go Bayesian, when bootstrap works perfectly fine.

[1] http://statweb.stanford.edu/~tibs/ElemStatLearn/ [2] http://www.wiley.com/WileyCDA/WileyTitle/productCd-111870860... [3] https://en.wikipedia.org/wiki/Conjugate_prior#Continuous_dis... [4] https://books.google.com/books/about/Time_Series_Analysis_by... [5] https://web.stanford.edu/~hastie/CASI/

Re: Frequentists should more often consider using Bayesian methods

#13
post #2

I'm getting frustrated by the Bayesian train at the moment, as its drawbacks get glossed over. "Oh yeah, there's priors, but they're not important for X, Y and Z reasons." In large samples, the frequentist and Bayesian methods are the same, so then why does it matter? In small samples, the prior becomes significant, so why use it if it shapes the estimates depending on what you use? You could turn these arguments for…

Frequencist methods do not remove the prior, they merely hide it, make it implicit. Usually the implicit prior is reasonable but being blind of your prior still risks making you trip on subtle trade-offs and biases for more complex problems.

[deleted]

Re: Frequentists should more often consider using Bayesian methods

#14
post #9
post #7

Earlier quoted context omitted.

If the inferential question you're interested in is, "given data X, what do I conclude about underlying cause/variable/parameter T?", then you are a Bayesian, like it or not.* Sure you can define a likelihood function p(X|T), but that doesn't give you p(T|X) unless you multiply by a prior p(T). Now certainly p(T|X) is not always the question, but in the vast majority of cases people do want to use data to draw conclu…

No, if you are drawing conclusions from only the data presented you are not doing Bayesian. Further, there are more than 2 options.

and they are?

Re: Frequentists should more often consider using Bayesian methods

#16
What are the Bayesian methods I would use if I have multiple related outcomes and multiple predictors...you know...the situation in which many scientists find themselves with longitudinal data and typically turn to repeated measures and random effect general linear models?

Re: Frequentists should more often consider using Bayesian methods

#17
post #11

Maybe it's because my formal math training is not in probability and statistics, but it's so bizarre to me that in a technical situation people would let a philosophical position dictate their approach rather than best tools for the job. Sometimes I'll solve a math problem analytically, and sometimes its easier to do it numerically. But it would be foolish for me to take a hardline stance on one vs the other. Rather…

How do you decide which tool is the best for the job?

Re: Frequentists should more often consider using Bayesian methods

#18
post #2

I'm getting frustrated by the Bayesian train at the moment, as its drawbacks get glossed over. "Oh yeah, there's priors, but they're not important for X, Y and Z reasons." In large samples, the frequentist and Bayesian methods are the same, so then why does it matter? In small samples, the prior becomes significant, so why use it if it shapes the estimates depending on what you use? You could turn these arguments for…

If you are trying to answer the same question in the same way there isn't much if a difference. Frequentist and bayesian statistics are different paradigms , not different number crunchers. --- I talk about this with graphs and such in my blog post ( https://www.lucidchart.com/blog/2016/10/20/the-fatal-flaw-of... ) but here's the short version: Common question in our line of work: when should I end my A/B test? End t…

Actually, in the frequentist paradigm you could choose to run a sequential hypothesis test which will end when you've acquired sufficient data[1]. Or, if you want to get fancy you could use a multi-armed bandit approach which is probably optimal in many situations in perhaps a more robust way than many Bayesian methods[2]. Really both can work well. My advice is, use whichever you know well enough to utilize effectively!

[1]: https://en.m.wikipedia.org/wiki/Sequential_analysis

[2]: https://en.m.wikipedia.org/wiki/Multi-armed_bandit

Re: Frequentists should more often consider using Bayesian methods

#19
post #9
post #7

Earlier quoted context omitted.

If the inferential question you're interested in is, "given data X, what do I conclude about underlying cause/variable/parameter T?", then you are a Bayesian, like it or not.* Sure you can define a likelihood function p(X|T), but that doesn't give you p(T|X) unless you multiply by a prior p(T). Now certainly p(T|X) is not always the question, but in the vast majority of cases people do want to use data to draw conclu…

No, if you are drawing conclusions from only the data presented you are not doing Bayesian. Further, there are more than 2 options.

you can draw conclusions from only the data presented to you doing bayesian inference.

Re: Frequentists should more often consider using Bayesian methods

#20
post #9
post #7

Earlier quoted context omitted.

If the inferential question you're interested in is, "given data X, what do I conclude about underlying cause/variable/parameter T?", then you are a Bayesian, like it or not.* Sure you can define a likelihood function p(X|T), but that doesn't give you p(T|X) unless you multiply by a prior p(T). Now certainly p(T|X) is not always the question, but in the vast majority of cases people do want to use data to draw conclu…

No, if you are drawing conclusions from only the data presented you are not doing Bayesian. Further, there are more than 2 options.

Have you heard of the universal prior? It allows you to do exactly that with Bayesian statistics.

Although it is so esoteric that one might argue that it is more the field of algorithmic statistics which merely employs Bayesian.

https://en.m.wikipedia.org/wiki/Algorithmic_probability https://en.m.wikipedia.org/wiki/Solomonoff%27s_theory_of_ind...

Post reply on HN