Earlier quoted context omitted.
They’ve added incorrect answers?
No; those were there already. I assumed they'd added the "does not double down" behaviour. Bing Chat (also an OpenAI model, but presumably with a different set of fine-tuning / filters to ChatGPT – possibly the bare OpenAI API) doubles down in situations like this: https://nitter.dark.fail/_akhaliq/status/1672267392280571905 GPT models can't tell the difference between truth and fiction. All you can choose by fine-tu…
But that’s beyond the point. The question is, how would you include incorrect responses into the training. In a way that it would not increase the probability of the model to give an incorrect response?
I guess you can maybe train with a mix of correct and incorrect responses, hallucinations and nonsense in the conversation, but then make clear that the responses were incorrect, adding context to these. And then fine-tune the AI actor to avoid giving incorrect responses or hallucinations altogether.