And Intel has big.LITTLE but they're not using TSMC nodes. Apple comparisons run on OSX or Asahi Linux, which doesn't run x86, so the software/ecosystem is different. Etc etc.
No comparison is ever perfect, you'll just have to take the best data we have and run with it. Complaining that a study isn't exactly perfect is trite, once you get beyond the "junior scientist makes obvious methodological error" tier it's honestly one of the least useful forms of criticism, someone is always going to think it should have been done better/differently (and wants you to take the time and spend the money to do it for them). But science is about doing the best you have and trying to make reasonable extrapolations about the things you can't.
Single-thread benchmarks on big vs little cores will get you IPC figures, and then you can scale those according to clocks you see on full-load conditions, for example.
Or you can simply compare it to a future 13th/14th gen Intel i3 with big.LITTLE, or an AMD quad-core APU. AMD has SMT, that's an advantage, Apple has single-threaded cores but a couple extra little cores, it's similar-ish.
Nothing is ever perfect. You just make do. It doesn't mean we throw up our hands and scream that if we can't be accurate to 1000 decimal points then we can never truly know anything.
(You may not have intended this, so just FYI: it kinda comes off like you're pre-stating that you won't accept the results if they don't come out the way you like, that you'll find some other difference between the two to latch onto. And there will always be some minor thing you can latch onto, no two designs are exactly identical. But that's not really an honest way to approach science, merely being able to theorize some differences isn't useful and if you feel strongly about it then you should do a similar test yourself to demonstrate.)