LLM-Generated Tests Influence Implementation Selection Outcomes
Key findings on the impact of LLM-generated tests on implementation selection and evaluation.
Large language models are increasingly utilized to generate tests and executable evidence for evaluating candidate code. This raises comparative measurement challenges when the evidence used for scoring is generated after observing the candidates. The study reveals that the visibility of candidates during test generation can shift evaluation outcomes, leading to significant winner reversals in certain cases.