» Tag
swe-bench
2 postsIf 30% of SWE-Bench Pro Tasks Are Broken, Add an Uncertainty Budget
OpenAI's SWE-Bench Pro audit found ~30% of tasks broken. Learn how to version task validity and report uncertainty intervals in benchmarks.
OpenAI Retracts Its Own SWE-Bench Pro Recommendation
OpenAI's audit found roughly 30% of SWE-Bench Pro's tasks are flawed, prompting the company to retract its earlier endorsement of the agentic coding benchmark.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com