Sorry to be a wet blanket, but:
When looking at studies, there are two types of validity to pay attention to, "external" and "internal." Internal is the sort that says: does this study actually test what it claims to? If a study looks in a black box, sees there's no elephant inside, and says "there's no elephant in there! Huzzah, we've proven it's a tiger!", that's an internally invalid study. External validity is about applicability: if I give a dose of abx to a bunch of people with advanced AIDS to treat an infection, and find it doesn't work, it doesn't mean "this abx doesn't work." It means "this abx doesn't work /in this population/. Don't extrapolate it to an immune-competent population."
1. The study begins with what they describe in their registration documents as "healthy elderly." That's a bit kind: "Sixty-two normal volunteers, who responded to a local advertisement, were screened. Subjects with any neurological condition, metallic implants, claustrophobia, tinnitus, BMI ≤30, high blood pressure (systolic≤140 mmHg), diabetes mellitus, intensive physical engagement (more than 1 hour/week) and abnormal performance in a cognitive screening test (MMSE 13) [33] were excluded." In short, they started with an anomalously healthy population, who likely have a lifetime of exposure to exercise and physical activity. Do the results here extrapolate to the general population? Do they extrapolate to the same degree? It's a decent question mark, considering their results are non-significant in everything except their "looks like false positives" brain volume measures to begin with - even a little bit of "ehh... maybe not so much" takes them into the realm of "no effects", alongside every other measure in the paper. External validity is doubtful here.
2. Internal validity is a bit doubtful too, first for reasons of selection bias. They recruited 62 healthy volunteers; they had 14 dropouts. A 22% dropout rate isn't atrocious, but it's more than enough - if it's not random - to skew a study. Their enrollment figure describes the dropouts as almost entirely pre-randomisation (10/14), but the study description notes that 6 drop-outs were due to failure to achieve frequency of adherence, which clearly had to occur post-randomization, and 2 due to dissatisfaction with group assignment (obviously post-randomization). 6 got seriously ill - I'd love to know in which group, and with what.
3. Their way of controlling for equivalent physical load was to measure heart rate twice. On the one hand, not crazy. On the other hand, if you've ever seen your HR during a workout session, you'll see how noisy that is - and with a sample of a whole 38 pairs, that's a relatively huge amount of noise. Moreover, it doesn't appear that they used that to guide intensity of intervention, just "to control" (which I take to mean, to plug into a multivariate model at some point - except they don't, as they describe the covariates they plug into their model to be age, sex, and intracranial volume for brain volume t-tests) I'm skeptical this is an adequate control - I'd at least have wanted a time-weighted average HR.
3.b. The sports intervention had one effort-controllable component (sport bike), but had 3 different components. It's not at all clear they could capture the effort under the regimen above, especially as the relatively light strength exercise that the elderly tend to tolerate is the place where I'm most suspicious of them failing to capture an effort delta.
4. The differences in brain volume swung in different directions in each group, without making an awful lot of sense (more right cerebellar development in standard exercise group?? So, asymmetrically, the less-coordination-demanding intervention showed more development of the primary coordination center of the brain?). But more generally, looking at supplementary table S3, note that dancing showed improvements in anterior and posterior white matter, and standard exercise in temporal and occipital. The brain isn't that cleanly delineated - to find such statistically strong effects in such broad brushstrokes sets off a red flag for me.
The fact that there was scattered growth *in both groups suggests we're looking at false positives. Not precisely a new problem for this study methodology. Their p-threshold of .001 is considered best-practices (aka, the bare minimum) for cluster-based adjustment of multiple testing in neuroimaging: so that's good. They don't report the p-values on their brain volume testing; table S3 simply notes "p5. Dance group had lower BDNF plasma at baseline vs. standard exercise (1500 vs. 2100) (p .14), and equal at post (2200 to 2100, p .6). In short, they showed regression to the mean in the dance group. They report this as "the dance group had an increase in plasma BDNF from baseline." Serum levels likewise were not significantly difference pre and post between the two groups (dance went from 35K to 36K, sport went from 30K to 29K). In short, nothing happened. But they hid the "fucking nothing happened" in supplementary table 4, and dressed it up real pretty in the included figure 4.
6. Cognitive outcomes, the only thing that actually matters here: no differences.
7. At least some physical fitness differences? No, none there either.
TL;DR they found nothing, made some misleading figures out of it. There are no perfect studies, but this one just boils down to "found nothing, needed publication."