Live data from Hacker News

Clinical trials launch to test coronavirus treatments

nature.com

11–20 of 33 posts

Re: Clinical trials launch to test coronavirus treatments

#11

The official infection and death numbers in China appear to be completely fabricated as they followed an almost perfect quadratic progression: https://old.reddit.com/r/dataisbeautiful/comments/ez13dv/oc_... None of the various measures taken to contain the outbreak have affected those numbers. This apparently has happened before with organ donation data: https://bmcmedethics.biomedcentral.com/articles/10.1186/s129...

This is a breathtakingly stupid abuse of data analysis. Which of course means it's at the top of Reddit's r/bestof, with thousands of upvotes.

Let's take a look at that paper you linked. The supposedly damning point is that you can fit the data extremely well with a quadratic, in the sense that R^2 is high. However, the data points are rapidly increasing, and because of the way R^2 works, that means only the last ~4 data points have any impact on R^2 at all. The statement of the paper boils down to the claim that you can fit 4 data points on a smooth curve accurately using a 3-parameter fit whose form you got to choose. This is obviously true, for any 4 data points.

It's too bad that reason goes out the window whenever China is involved.

Re: Clinical trials launch to test coronavirus treatments

#13
post #11

The official infection and death numbers in China appear to be completely fabricated as they followed an almost perfect quadratic progression: https://old.reddit.com/r/dataisbeautiful/comments/ez13dv/oc_... None of the various measures taken to contain the outbreak have affected those numbers. This apparently has happened before with organ donation data: https://bmcmedethics.biomedcentral.com/articles/10.1186/s129...

This is a breathtakingly stupid abuse of data analysis. Which of course means it's at the top of Reddit's r/bestof, with thousands of upvotes. Let's take a look at that paper you linked. The supposedly damning point is that you can fit the data extremely well with a quadratic, in the sense that R^2 is high. However, the data points are rapidly increasing, and because of the way R^2 works, that means only the last ~4…

I'm sorry, but I don't quite understand your point. There are far more than 4 data points fitting the curve smoothly. Shouldn't the data have a lot more variation for something like this?

Re: Clinical trials launch to test coronavirus treatments

#14
post #9

The official infection and death numbers in China appear to be completely fabricated as they followed an almost perfect quadratic progression: https://old.reddit.com/r/dataisbeautiful/comments/ez13dv/oc_... None of the various measures taken to contain the outbreak have affected those numbers. This apparently has happened before with organ donation data: https://bmcmedethics.biomedcentral.com/articles/10.1186/s129...

For all you know, this is the best-case outcome with countermeasures.

I would have expected the rolling out of various countermeasures as the epidemic spreads to have a visible effects on the figures. It's weird that it fits a predictable progression so smoothly.

Re: Clinical trials launch to test coronavirus treatments

#15
post #2

This is crazy fast. Usually hit to target takes almost a year, and they are doing this all in less than a month.

Quite sure they had a Zika virus vaccine completed within a few months after the outbreak. Obviously getting regulatory approval takes much, much longer but people would probably be surprised how quickly such things can be developed.

Re: Clinical trials launch to test coronavirus treatments

#16
post #11

Earlier quoted context omitted.

This is a breathtakingly stupid abuse of data analysis. Which of course means it's at the top of Reddit's r/bestof, with thousands of upvotes. Let's take a look at that paper you linked. The supposedly damning point is that you can fit the data extremely well with a quadratic, in the sense that R^2 is high. However, the data points are rapidly increasing, and because of the way R^2 works, that means only the last ~4…

I'm sorry, but I don't quite understand your point. There are far more than 4 data points fitting the curve smoothly. Shouldn't the data have a lot more variation for something like this?

I'm responding to the paper you linked, where you assert "this has happened before".

The story with the Reddit link is similar: at the time of the posting, only ~6 data points actually mattered for R^2, so you would expect R^2 to be very high. (The fit is extremely poor for the first ~6 data points, but that doesn't affect R^2 at all.) After the last date on the chart shown, the R^2 begins to plummet. This is also exactly what you would generically expect.

Amusingly, both of these totally ordinary results are taken as evidence for a conspiracy: the 6 points that fit are taken to be fabricated, and the subsequent points are taken to be fabricated with the sole purpose of throwing off people on Reddit. But if you can claim evidence for a conspiracy no matter what the numbers are, has something gone wrong in your reasoning process? You be the judge!

Re: Clinical trials launch to test coronavirus treatments

#17

Interesting ethics of using people as the null hypothesis population knowing they are then condemned to death

The alternative is to treat everyone as the null hypothesis population, or administer the treatment to everyone and then be unable to determine if it's effective.

Both of these are strictly worse, since neither of them can result in the discovery of an effective treatment.

Re: Clinical trials launch to test coronavirus treatments

#19
post #16

Earlier quoted context omitted.

I'm sorry, but I don't quite understand your point. There are far more than 4 data points fitting the curve smoothly. Shouldn't the data have a lot more variation for something like this?

I'm responding to the paper you linked, where you assert "this has happened before". The story with the Reddit link is similar: at the time of the posting, only ~6 data points actually mattered for R^2, so you would expect R^2 to be very high. (The fit is extremely poor for the first ~6 data points, but that doesn't affect R^2 at all.) After the last date on the chart shown, the R^2 begins to plummet. This is also ex…

I would say the numbers deviating from the model after it has been publicized and discussed online are still evidence against the data being fabricated, but weaker than it could have been because there were definitely Chinese people watching those discussion on social media, and so the idea that they might have corrected the mistake and started adding more variation isn't that unlikely.
Post reply on HN