» Tag
llm-evaluation
3 posts«August 2026
Audit Finds LLM Agent Credit Signals No Better Than Random Chance
A causal audit in ALFWorld shows judge- and confidence-based credit signals for training LLM agents perform no better than chance at identifying key steps.
SciCode-Verified: Flawed Benchmark Masked True LLM Coding Skill
Fixing 263 defects in the SciCode benchmark lifted LLM scientific-coding accuracy from 60% to 98%, revealing the test—not the models—was flawed.
data-eng-bench: Snowflake's dbt benchmark for coding agents
Snowflake's data-eng-bench tests coding agents on 103 realistic dbt data-engineering tasks across DuckDB and Snowflake, via the open-source Harbor framework.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com