> If this was an evolution in the wild, then thousands, maybe millions, of "almost SARS CoV2" cases should surface after years of search
I question that assumption. Evolution happens in tiny bursts (founder organisms, population bottlenecks, etc) in between which are relatively long periods of Hardy-Weinberg equilibrium (i.e. periods where evolution is not happening). That's why when we look into the fossil record we're (typically) able classify our findings as one of several stable species, rather than some sort of smooth spectrum in which everything is equally represented a transitional state to something else.
There's a very strong sampling bias where you're much more likely to find an organism in equilibrium and not an individual that was part of whatever transition you're interested in. If this were not the case, the existence of evolution would've been a no-brainer rather than a discovery that biologists had to fight to establish.
In the last 4 years there have been 20 variants of concern, but currently there are only two (https://www.ecdc.europa.eu/en/covid-19/variants-concern). So on average these populations only last something like six months before they're outcompeted by a new variant--and that's six months for populations in equilibrium. If you're looking to collect a sample in a non-equilibrium state, it seems likely to me that your window for doing so might only last a few weeks.
So they're looking for the missing link, and they haven't found it yet. But it's not clear how likely they are to ever find it, so the jump from "they haven't found it" to "it was never in nature" is not warranted.
One experiment that could be done, which I think would shed light on the situation, would be sample modern humans in an attempt to find the first variant (discovered in 2019). If we find it easily, maybe we should expect to find its ancestors relatively easily, supposing they're out there (never mind that animals are harder to work with than humans). Then again, it may be extinct, in which case we should probably not hold our breaths about finding its ancestor either.
But sequencing experiments cost on the order of $100-$1000 to do, so neither my proposed experiment nor the quest to find the reservoir species are likely to happen at a scale of
> thousands, maybe millions, of ... cases
Sure, you might get more than one virus per sample, but typically the way to sort that out is to align the reads (typically 100-200 nucleotides long) to a reference genome. It works well when your sample came from one individual with genetically identical cells. I'm not sure how quickly things get murky when you've got a pile of reads from a sample with hundreds or thousands of genetically dissimilar viral particles, plus whatever non-viral DNA ended up in the sample, but I'd be surprised if you could get more than 10 or 20 reliable alignments out of it. Eventually your error bars are going to get too big and you're not going to know which data goes which which sites. I haven't worked with an assay like that before but I'm positive that there are going to be limits to how much you can squeeze out of a single run.