179 questions
No questions match those filters.
GPT-4-as-judge is widely used but has known biases. Nam...
This is one of the questions in the full AI/ML interview bank. Pro unlocks all 1789 questions; Premium includes the same bank plus the highest daily Practice limit.
See plansGPT-4-as-judge is a common evaluation technique, but it suffers from specific biases. Position bias occurs when the judge favors responses in the first position; this is mitigated by randomizing the order of responses and averaging the outcomes. Verbosity bias occurs when the judge favors longer responses regardless of quality; this is mitigated by constraining response lengths or instructing the judge to penalize unnecessary verbosity. Other considerations include self-enhancement bias, where the judge favors outputs similar to its own training distribution, which can be addressed by using multiple judge models or calibrating against human judgments.