Live data from Hacker News

20 lines of code that beat A/B testing (2012)

stevehanov.ca

151–160 of 164 posts

Re: 20 lines of code that beat A/B testing (2012)

#151
post #124

Earlier quoted context omitted.

that's an appeal to authority, sounds like @kens actually tried to verify this story. I'd be curious to hear more examples as well. I believe I heard something similar on You're Not So Smart podcast but that might have referenced the same example

I do think the example is bullshit, simply because you could throw ONE very thick pot and have it weight more than 50 pounds. Of course it's a horrible pot and you wouldn't get any better since you only made one piece, but it would get you an A since it would use more clay. Total # of pieces thrown would make more sense than pounds of clay used.

Forgive me if I've missed something but the original comment says the group that threw the most clay scored higher on creativity, not scored higher on pounds of clay thrown.

Re: 20 lines of code that beat A/B testing (2012)

#152
post #76

Here's what everyone is missing. Don't use bandits to A/B test UI elements, use them to optimize your content / mobile game levels. My app, 7 Second Meditation, is solid 5 stars, 100+ reviews because I use bandits to optimize my content. By having the system automatically separate the wheat from the chaff, I am free to just spew out content regardless of its quality. This allows me to let go of perfectionism and just…

You might want to be careful with your conclusions. We don't know from this study whether the second group was more or less likely to get trapped in the Expert Beginner phase of development. You definitely don't get anywhere without practice, but you are likely to get nowhere fast without theory.

insightful analogy.

so then "Expert Beginner phase" = local minima, which might be reached quickly compared to a more systematic search (practice informed/infused with theory) and once in that local state, the "expert beginner" might be unable to escape because of over-reliance on small, poorly guided, step size; the real prize anyway is the global minima.

Re: 20 lines of code that beat A/B testing (2012)

#153
post #111
post #101

Earlier quoted context omitted.

I tried to find the original source of the quality-vs-quantity pottery class story a while back. I think it originates in the book "Art and Fear" but in that book it reads like a parable rather than a factual event. I'm highly suspicious of whether this event actually happened. Anyone have solid evidence?

It was featured in "Thinking Fast and Slow", and those authors seem quite academically rigorous.

Yes they're very good at "seeming" the part. Until you check the cites/references.

Please excuse me for lumping Kahneman, Tversky and Taleb into the same bin (stating beforehand, feel free to dismiss or form your own opinions because of this). My justification is they cite eachother all the time, write about the same topics and are quoted as doing their research by "bouncing ideas off each other" (only to later dig up surveys or concoct experiments to confirm these--psych & social sciences should take some cues from psi research and do pre-registration).

This is now the second time I notice that one of their anecdotes posed as "research" doesn't quite check out to be as juicy and clear-cut (or even describing the same thing) as the citation given.

The other one was the first random cite I checked in Taleb's Black Swan. A statistics puzzle about big and a smaller hospital and the number of baby boys born in them on a specific day. The claim being research (from a metastudy by Kahneman & Tversky) showing a large pct of professional statisticians getting the answer to this puzzle wrong. Which is quite hard to believe because the puzzle isn't very tricky at all. Checking the cited publication (and the actual survey it meta'd about), turns out it was a much harder statistical question, and it wasn't professional statisticians but psychologists at a conference getting the answer wrong.

Good thing I already finished the book before checking the cite. I was so put off by this finding, I didn't bother to check anything else (most papers I read are on computational science, where this kind of fudging of details is neither worth it, nor easy to get away with).

Which is too bad because the topics they write about are very interesting, and worthwhile/important areas of research. I still believe it's not unlikely that a lot of their ideas do in fact hold kernels of truth, but I'm fairly sure that also a number of them really do not. And their field of research would be stronger for it to figure out which is which. Unfortunately this is not always the case for the juicy anecdotes that sell pop-sci / pop-psych books.

Re: 20 lines of code that beat A/B testing (2012)

#154
post #111

Earlier quoted context omitted.

It was featured in "Thinking Fast and Slow", and those authors seem quite academically rigorous.

Yes they're very good at "seeming" the part. Until you check the cites/references. Please excuse me for lumping Kahneman, Tversky and Taleb into the same bin (stating beforehand, feel free to dismiss or form your own opinions because of this). My justification is they cite eachother all the time, write about the same topics and are quoted as doing their research by "bouncing ideas off each other" (only to later dig u…

This is very interesting, thank you. Modern non-fiction is so frustrating.

Re: 20 lines of code that beat A/B testing (2012)

#155
post #76

Here's what everyone is missing. Don't use bandits to A/B test UI elements, use them to optimize your content / mobile game levels. My app, 7 Second Meditation, is solid 5 stars, 100+ reviews because I use bandits to optimize my content. By having the system automatically separate the wheat from the chaff, I am free to just spew out content regardless of its quality. This allows me to let go of perfectionism and just…

solid 5 stars, 100+ reviews because I use bandits to optimize my content.

It's great app, but were the ratings lower in the beginning before the optimization or how do you know the optimization helped the ratings? I'm asking because it seems the app could have good ratings regardless of the content optimization because it's a "feel good" app. Are there counterexamples of of other meditation apps were the UI is good but it has bad reviews because of their low quality content?

Re: 20 lines of code that beat A/B testing (2012)

#156
post #124

Earlier quoted context omitted.

I do think the example is bullshit, simply because you could throw ONE very thick pot and have it weight more than 50 pounds. Of course it's a horrible pot and you wouldn't get any better since you only made one piece, but it would get you an A since it would use more clay. Total # of pieces thrown would make more sense than pounds of clay used.

Forgive me if I've missed something but the original comment says the group that threw the most clay scored higher on creativity, not scored higher on pounds of clay thrown.

Guess you missed something. "The second group was graded on only the total number of pounds of clay they threw."

Re: 20 lines of code that beat A/B testing (2012)

#157
post #124

Earlier quoted context omitted.

that's an appeal to authority, sounds like @kens actually tried to verify this story. I'd be curious to hear more examples as well. I believe I heard something similar on You're Not So Smart podcast but that might have referenced the same example

I do think the example is bullshit, simply because you could throw ONE very thick pot and have it weight more than 50 pounds. Of course it's a horrible pot and you wouldn't get any better since you only made one piece, but it would get you an A since it would use more clay. Total # of pieces thrown would make more sense than pounds of clay used.

(Not that I think the anecdote is real, but) that's the idea -- the parable was about how the pot with the best quality was found in the 'quantity' team even though this wasn't their objective.

Re: 20 lines of code that beat A/B testing (2012)

#158
post #111

Earlier quoted context omitted.

It was featured in "Thinking Fast and Slow", and those authors seem quite academically rigorous.

Yes they're very good at "seeming" the part. Until you check the cites/references. Please excuse me for lumping Kahneman, Tversky and Taleb into the same bin (stating beforehand, feel free to dismiss or form your own opinions because of this). My justification is they cite eachother all the time, write about the same topics and are quoted as doing their research by "bouncing ideas off each other" (only to later dig u…

You should read silent risk (currently freely available as a draft on talebs website). It's mathematical, with solid proofs and derivations, so it doesn't really suffer from the same problems as The Black Swan. I had basically the same issues with the Black Swan that you did.

Re: 20 lines of code that beat A/B testing (2012)

#159

Earlier quoted context omitted.

That's not thoroughly ridiculous. For many problems in machine learning, k-nearest-neighbors and a large dataset is very hard to beat in terms of error rate. Of course, the time to run a query is beyond atrocious, so other models are favored even if kNN has a lower error rate.

According to [1], k-NN is pretty far from being the top general classifier. Admittedly, all the data sets are small, with no more than 130,000 instances. When does it start becoming "very hard to beat", and what are you basing that on? 1. http://jmlr.org/papers/volume15/delgado14a/delgado14a.pdf

The vast majority of the data sets are very small (on the order of 100 to 1000 examples!). In fact, in the paper they discarded some of the larger UCI datasets.

It's not surprising that they found random forests perform better, even though the conventional wisdom is boosting outperforms (it's much harder to overfit with forests).

Re: 20 lines of code that beat A/B testing (2012)

#160

Earlier quoted context omitted.

> but they never really published much justifying it Not that I'm trying to defend Optimizely (I'm not a huge fan, but for other reasons...). I can't vouch for the quality either, but they did publish something about it[0] - that at least looks quite scientific. Happy to read any critique of course. [0] http://pages.optimizely.com/rs/optimizely/images/stats_engin...

Latex is a wonderful way to make a marketing paper look like a scientific one. It doesn't accurately describe the method, but that isn't really its purpose. It's a more technical description of the blog post, meant for people using the product to understand some of the tradeoffs and get more accurate results. They are still having people make very fundamentally flawed assumptions about the data, which results in inco…

Hi, this is Leo, Optimizely's statistician. If you're looking for a more scientific paper, maybe take a look at this one we wrote recently: http://arxiv.org/abs/1512.04922

Should have everything you would ever want to know about the method.

I agree with you that the problem of inference and interpretation between A/B data, algorithms, and the people who make decisions from them is a hard one and worth working on.

That said, I do think the two sources of error our stats engine addresses - repeatedly checking results, and cherry picking from many metrics and variations - did make progress in having folks correctly interpret A/B Tests. This did result in more conservative results, but the benefit was that the variations that do become significant are more trustworthy. I think this was absolutely the right tradeoff to make for our customers, and trustworthyness is a pretty important aspiration for stats/ML/data science in general.

Of course I did write the thing, so I'm not very impartial.

Post reply on HN