The three biases that ruin judges
Verbosity bias: judges prefer longer answers regardless of quality. Position bias: in pairwise comparison they favour whichever answer came first. Self-preference: a model rates its own outputs above others.
The generated prompt states explicitly that length is not quality and that answer order is random. Neither eliminates the bias, but both measurably reduce it.
The swap test
The pairwise output ends with the instruction to run the comparison twice with the answers swapped. If the verdict flips, the judge is following position rather than quality and the result should be discarded.
It doubles your judging cost and it is the difference between an evaluation you can act on and a number that feels like one. Run it at least while you are calibrating.