Live data from Hacker News

Understanding Stein's Paradox (2021)

joe-antognini.github.io

1–10 of 50 posts

Re: Understanding Stein's Paradox (2021)

#3
Sorry, I'm siding with the physicists here. If you're going to declare that your seemingly arbitrary choice of coordinate system is actually not arbitrary and part of your prior information about where the mean of the distribution is suspected to be, you have to put that in the initial problem statement.

Re: Understanding Stein's Paradox (2021)

#4
I do not get it: if variance is too large, a random sample is very little representative of the mean. As simple as that?

Now the specific formula may be complicated. But otherwise I do not understand the “paradox”? Or am I missing something?

Re: Understanding Stein's Paradox (2021)

#5

I'm horrible at stats, but is this saying that if I have 5 jars of pennies, and I guess the amount in each one. Then I find the average of all my guesses, and the variance between the guesses, then I can adjust each guess to a more likely answer with this method?

No, I don't think these problems are related.

Re: Understanding Stein's Paradox (2021)

#6

Sorry, I'm siding with the physicists here. If you're going to declare that your seemingly arbitrary choice of coordinate system is actually not arbitrary and part of your prior information about where the mean of the distribution is suspected to be, you have to put that in the initial problem statement.

There is nothing magical about the origin, the shrinkage can be done towards any point and in fact when estimating multiple means it's customary to move each point closer to their average.

https://www.math.drexel.edu/~tolya/EfronMorris.pdf

Re: Understanding Stein's Paradox (2021)

#7

I'm horrible at stats, but is this saying that if I have 5 jars of pennies, and I guess the amount in each one. Then I find the average of all my guesses, and the variance between the guesses, then I can adjust each guess to a more likely answer with this method?

Not necessarily "more likely" but "better" in some "loss" sense.

It could be "more likely" in the jars example where estimates may convey some relevant information for each other. But consider this example from wikipedia:

"Suppose we are to estimate three unrelated parameters, such as the US wheat yield for 1993, the number of spectators at the Wimbledon tennis tournament in 2001, and the weight of a randomly chosen candy bar from the supermarket. Suppose we have independent Gaussian measurements of each of these quantities. Stein's example now tells us that we can get a better estimate (on average) for the vector of three parameters by simultaneously using the three unrelated measurements."

https://en.wikipedia.org/wiki/Stein%27s_example#Example

Re: Understanding Stein's Paradox (2021)

#8

Sorry, I'm siding with the physicists here. If you're going to declare that your seemingly arbitrary choice of coordinate system is actually not arbitrary and part of your prior information about where the mean of the distribution is suspected to be, you have to put that in the initial problem statement.

You can put the origin anywhere and for almost all choices the adjustment is almost zero. But if the choice happens to be very close to the sample point, against all (prior) probabilities, then that fact affects the prior.

Re: Understanding Stein's Paradox (2021)

#9
post #4

I do not get it: if variance is too large, a random sample is very little representative of the mean. As simple as that? Now the specific formula may be complicated. But otherwise I do not understand the “paradox”? Or am I missing something?

In the 1D case the single point will be the best estimator for the mean no matter what the variance is

Re: Understanding Stein's Paradox (2021)

#10
Stein's paradox is bogus. Somebody needs to say that.

Here's one wikipedia example:

  > Suppose we are to estimate three unrelated parameters, such as the US wheat yield for 1993, the number of spectators at the Wimbledon tennis tournament in 2001, and the weight of a randomly chosen candy bar from the supermarket. Suppose we have independent Gaussian measurements of each of these quantities. Stein's example now tells us that we can get a better estimate (on average) for the vector of three parameters by simultaneously using the three unrelated measurements. 
Here's what's bogus about this: the "better estimate (on average)" is mathematically true ... for a certain definition of "better estimate". But whatever that definition is, it is irrelevant to the real world. If you believe you get a better estimate of the US wheat yield by estimating also the number of Wimbledon spectators and the weight of a candy bar in a shop, then you probably believe in telepathy and astrology too.
Post reply on HN