Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation
This study explores how prior scores in LLM-as-a-Judge systems affect evaluations, emphasizing the need for careful context engineering.
Large language models (LLMs) are increasingly used to evaluate generated content, leading to the LLM-as-a-Judge paradigm. This study tests the assumption that evaluations are independent of prior scores, revealing that previous scores can anchor judgments and skew ratings. The findings highlight the need for careful context engineering in LLM evaluations to ensure reliability and mitigate bias.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work