@tokalator4.2.0…installs…downloads
Dictionary
Evaluation

LLM-as-a-Judge

Using one AI to grade another AI's answers, like a teacher marking essays, which is fast but has its own biases.

Definition

Using a strong language model to grade or compare other models' outputs when no exact-match answer exists. Zheng et al. (2023) found GPT-4 judges agreed with human preferences over 80% of the time, about the same level at which humans agree with each other, but documented position, verbosity, and self-enhancement biases and limited reasoning ability.

Example

A judge prompt asks 'Which code review is more accurate, A or B?' and is rerun with the order swapped to cancel position bias.