» Tag
ai-benchmarks
1 postsSDABench: A New Benchmark Testing LLMs on Scientific Discovery
SDABench evaluates LLMs on six scientific capabilities beyond code execution, exposing major gaps in assumption selection and mechanistic reasoning.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com