Earlier quoted context omitted.
I expect that this will be tried, and I worry about some negative consequences if it works well. It could be a way of generating very effective propaganda, that defeats the efforts of the opposing LLM to call bullshit. On the other hand, diffusion models seem to have replaced GANs for image synthesis, so perhaps there's something I'm missing, or perhaps there's a way to combine both techniques.
I am not sure calling bullshit on images works the same way, given we generally want to invite creativity to image generation. Although an adversarial extra-finger detector seems in order! Extra fingers are not creativity. They are taboo!! But for text, there is a lot of structure around bullshitting, and fortunately, the internet is full of examples of people calling bullshit. As long as the bullshit adversary has t…
I'm not worried about using approaches like this to converge on truth (model A generates an answer to a question, model B attempts to refute model A). Perhaps we could train model A to generate lies that model B can't detect.