Have two calculus teachers, A and B, each with 20 students. Look at the final exam numbers. Put all the numbers in a bucket, stir briskly, draw out 20 numbers (test scores) and average, average the other 20, and get the difference in the averages. Do this many times. Get the empirical distribution of the differences in the averages.
Now look at the difference in the actual average for A and B. If this is out in a tail of the empirical distribution, then we reject the null hypothesis that the two teachers are equally good.
For this non-parametric, distribution-free, resampling, two-sample test, to make theorems about it, which should, will likely need at least an independence assumption and likely an i.i.d. (independent, identically distributed) assumption. Else maybe each student of teacher B is an older sibling of a student of teacher A!!!!
In a nutshell, what's going on in statistical hypothesis testing is that we make the null hypothesis, and that gives us enough assumptions, e.g., i.i.d., to calculate the probability of our calculated test statistic, e.g., the difference in the two averages, being way out in a tail. Without some such null hypothesis assumptions, we have no basis on which to reject anything, are not testing anything.
There's chance of getting all twisted out of shape philosophically( here: E.g., I outlined one statistical hypothesis test for the two teachers A and B. Okay, now consider ALL reasonably relevant* hypothesis: Maybe we on the test I outlined, the two teachers look very different, not equal, maybe teacher B better. But in ALL those hypothesis tests, maybe in one of the tests the two teachers look equally good or even teacher A looks better. Now what do we do? That is, there is a suspicion that teacher B looked better ONLY because of the particular test we chose. Maybe there has been some research to clean up this issue.