Vero: Can AI Agents Build Formally Verified Software Repositories?
Vero benchmarks AI agents' capabilities in joint implementation and proof synthesis for software repositories.
AI agents are increasingly utilized in software development but lack guarantees on the correctness of generated code. Vero introduces the first benchmark to assess coherent implementation and proof synthesis across multi-module codebases. With 43 instances from real-world repositories, Vero aims to enhance the reliability of AI-generated software.