GPT-4 definitely seems to be doing better on a lot of benchmarks and that's impressive. But it still hallucinates facts and I don't think anyone really has a good understanding of when and how that happens. Given that, is it really a good idea to be positioning this model as some kind of factual authority figure?
Human teachers make mistakes too - perhaps even at higher rates than GPT in some cases. Instead of isolating a single factor of hallucinations I think you have to also consider the cost and experience quality and make a more holistic judgement on whether these models are good or bad.
Just because humans make mistakes (or even as many or more mistakes than machines) doesn't mean the nature or consequences of the mistakes are the same.