In retrospect it wasn't really that the models were that bad compared to benchmarks, it was that Meta didn't work with inference engine providers ahead of time to integrate the new architecture like other teams usually do, so it was even more bugged than expected for the first few weeks.
The second problem that compounded it was that both L4 models were too large for 99.9% of people to run, so the only opinion most people had of it was what the few that could load it were saying after release. And they weren't saying good things.
So after inference was fixed the reputation stuck because hardly anyone can even run these behemoths to see otherwise. Meta really messed up in all ways they possibly could, short of releasing the wrong checkpoint or something.