Earlier quoted context omitted.
I don’t think they do. At least not in a safe way. The AI doesn’t understand meaning or the underlying material and concepts. I’m sure filters will be put into a pipeline but it will never be as good as letting the human sniff out the correct answers.
Today I asked chatGPT what Einstein's favourite food was. It gave a reasonable answer about him liking simple foods. The worrying thing is that I was satisfied with the result. It was plausible. It could well be true. There is a good chance that this kind of AI might provide the right kind of response that a large percentage of the public find convincing enough to not bother with additional research.
Example. People are persuaded that women can do two things at the same time and men cannot. Despite countless counterexamples, the original study positioning everyone on a bell curve with only 6% difference in time of execution, and without checking for the quality of the results (“254 plus 786 equals 126 quick mafs!!!”), it is blatantly false, but any search engine that would return that it is false would make itself rejected by humans.