DictionaryEvaluation
LLM-as-a-Judge
Using one AI to grade another AI's answers, like a teacher marking essays, which is fast but has its own biases.
Definition
Using a strong language model to grade or compare other models' outputs when no exact-match answer exists. Zheng et al. (2023) found GPT-4 judges agreed with human preferences over 80% of the time, about the same level at which humans agree with each other, but documented position, verbosity, and self-enhancement biases and limited reasoning ability.
Example
A judge prompt asks 'Which code review is more accurate, A or B?' and is rerun with the order swapped to cancel position bias.