Earlier quoted context omitted.
If your sample is blinded yet you consistently guess correctly, then presumably either 1. you failed at blinding or 2. there is a strong and discernible effect, regardless of what your other metrics might say (your other metrics could always be flawed after all). > no documentation of whether the events during the hour of test time were more or less stressful than those before it, and no taking the time of day, diet…
He has 94 data points. Not nearly enough to average out so many potentially confounding variables, and there is no way to know they would. That will be the case in almost all N=1 experiments. Perhaps he takes the pill at the onset of stress, and stress almost always tends to build afterwards. This would be a probable case for many people trying to use theanine in this way. I could never accept that we should presume…
That would not be a problem regarding the averaging I referred to, although it could well pose a problem for measurement depending on how it interacted with the selected metrics.
Note that the averaging I refer to is not regarding all possible values of some metric, but rather any discrepancy in the distribution of metrics which we expected to follow the same distribution between the sample and the control.
I think maybe there's a misunderstanding? It seems that we both agree that a variety of additional variable should be logged. I was not suggesting to omit them, but rather to use discrepancies in them to detect fundamental issues with the data or study design. I would also expect larger studies to do the same where possible.
At 94 data points it is entirely possible that there would be outliers that would have averaged out for a larger N but did not. In such a scenario the presence of such outliers should then be taken to indicate a problem with the data (ie the more discrepancies you observe, the less you should trust the data).