Author here. Methodology upfront because I'd ask the same things: Data: daily records from wearable users who logged sauna sessions via connected apps. Within-person design — each user is their own control, comparing their own sauna-day nights against their own non-sauna-day nights. No cross-user comparisons. Stats: paired t-tests, FDR-corrected p 0.2 threshold for "meaningful effect." Anything below d=0.2 we don't r…
- Is the wearable accurate enough to be sure that 3bpm is not a measurement fluke? - Why did you use the minimum heart rate value (which could be a measurement glitch) and did not compare a percentile (e.g., 2.5th lowest percentile)? - Were all assumptions for paired t-testing valid? How did you account for likely temporal correlations in the data (e.g., sauna could have an effect also on a night 2 days after it, same for exercise)? - How can you define a "comparable-intensity exercise day" if you don't know the characteristics of the sauna?