This is a good hypothetical. I like it. The "easy" win in such a procedure is to decide which debater is self-consistent. Unfortunately a consistent argument based on faulty premises yields false conclusions. However, since we're allowed to look at past debates for context, I think we can sidestep that issue. We need to determine which past debate-winning arguments led to objectively advantageous actions. This is a non-trivial problem. Let's additionally sidestep it by allowing the human being to weight outcomes based on past arguments.
Aside: Are we defining a 'settled' debate as one where the resultant action came from one side of the debate over the other, or are we additionally taking into account whether we consider the outcome of said action is positive or negative?
I think that it would be simple to identify logical arguments (well, fallacies anyway) which imply that the position being asserted is actually untenable. Identifying arguments which satisfy the inverse - that is, that the style of argument correlates with an advantageous action as a result, falls back into the trap where false premises lead to false conclusions.
Aside: I wonder if there are arguments where objectively false conclusions still lead to advantageous actions? Now there's an interesting question.
I suppose that if you could semantically dissect an argument and then correlate its structure against weighted outcomes of previous arguments this system could yield intriguing results. I would very much like to see what logical constructions correlate with poor outcomes. Perhaps it would reveal self-defeating human tendencies, or arguments which push aside rational behavior?
Then again, the strength of the Bayesian filter approach is that it is dumb. It doesn't have to do a semantic analysis. So we avoid going down the path of strong-ai or special-cased rules when using it. Does the purely textual content of a debate correlate strongly enough with the semantic content to distinguish between good arguments and bogus ones? I think the only way to tell would be to try.
I suspect that many of the most vigorous debates would end up revealing that the relative advantagion or bogosity of a given argument correlate most strongly with what your individual desired outcome is. Then we've just led ourselves to the question of which outcome is superior to the other.
Thanks for an interesting few moments of introspection!