Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure
The reliability of language model judges was audited, revealing significant discrepancies in repeatability.
Language model judges are now gatekeepers of training data, scoring outputs and driving leaderboards. However, they operate under the assumption that the same request to the same model will yield the same result tomorrow. We audited this assumption in two preregistered studies, finding significant discrepancies in repeatability across 52,988 attempts, which raises concerns about the reliability of these measurement instruments. This has implications for engineers regarding the stability of model evaluations.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work