When LLM Judges Agree: A Question of Trustworthiness
The independence of LLM judges is more important than majority voting. A new method aims to enhance reliability by modeling judge relationships.
Evaluating a retrieval-augmented-generation system involves a user asking a question, the system retrieving a text passage, and an LLM judge determining its relevance. However, the independence of judges is more critical than the total vote count. If judges make similar errors, the majority may not be as reliable as it seems. A new method addresses this issue by modeling the relationships between judges' outputs to achieve more trustworthy results.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work