> Any such model ought to have been independently reviewed before it is ever used for real policy decisions. Policy analysis is awash in models but no one ever really checks them. Going forward, health policy makers should ask for and disclose independent validation of any model before using its results to make recommendations of any consequence.
That's ignoring the time-limited nature of a virus response. "No decision" is itself a decision. Delay has a real cost, and using a potentially imperfect model is simply using the best information at hand.
I think everyone would agree that the ideal is to have well-documented, thoroughly-tested, easy-to-use, and bulletproof models to inform public response to every emergency. However, those models can't be built instantly, and that kind of bulletproofing is not directly relevant to the day-to-day research work aided by such models.
Adding research capacity to ensure that governments have a stable of well-researched, thoroughly-vetted models for emergencies would be a great thing, but it would also be quite expensive. Keeping any single model up to spec might be the job of 1FTE (so an additional ~$150k/grant/yr -- large for a research budget but small for a government), but that would have to be multiplied by every area where the government might possibly want research-informed decisionmaking on short notice.
> Even if the documentation, coding, and testing problems were fixed, the model logic is fatally flawed, which is evidenced by its poor forecasting performance.
That sounds like a great scientific criticism! Model validation is the cornerstone of research on computer models of things, and finding a poor forecast opens the door to many great research questions.
But without that further research, "poor performance" sounds a (loud) note of caution, but isn't necessarily fatal. The leading-order problem is to understand why the model performed poorly: was it improperly calibrated with information known at the time (such as if the virus behaves differently than assumed)? Was there some out-of-sample feature of the forecast that the model would not expect to do well on (e.g., low death rates in an open society because everything was shut down for severe weather anyway?) Is an overall trend correct but the timing in error?
"Flawed," I think, can be easily shown, and this should probably be expected in a research model. "Fatally flawed," however, is a stronger claim that must pass a greater burden of proof.