» Tag
benchmarks
35 postsTutorMoments: Testing If AI Tutors Know When to Help or Hold Back
Allen AI's TutorMoments benchmark tests whether LLM tutors know when to scaffold and when to push students toward harder reasoning.
Generating the database from the answer key in text-to-SQL benchmarks
A UIUC audit found over half of BIRD and Spider 2.0 annotations wrong. One developer inverted the process: declare the answer first, then generate a database that satisfies it.
Distilling DeepSeek into GPT-OSS Doesn't Transfer Its Censorship
Research shows distilling DeepSeek into GPT-OSS-120B boosts financial reasoning without inheriting the teacher model's political censorship.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comAMD Epyc 9996 'Venice': 256-core Zen 6 chip claims 3.4x lead over Xeon
AMD's 256-core Epyc 9996 Venice Zen 6 chip claims up to 3.4x throughput over Intel Xeon 6980P and 2.2x over Nvidia Vera in early benchmarks.
AI Coding Agents Are Absorbing the Leaves, Not the Whole Tree
SWE-bench and METR data track fast AI coding gains, but verification cost — not raw difficulty — still defines where agents stop.
WANDR Benchmark Tests AI Agents on Wide-and-Deep Research Tasks
WANDR benchmark evaluates AI research agents on wide-and-deep data collection tasks using reference-free, evidence-verified grading across 500 tasks.
Text-to-SQL Fails Because of Your Warehouse, Not the Model
Text-to-SQL accuracy collapses on real enterprise warehouses despite strong benchmark scores. Evidence points to undocumented schemas, not model quality, as the cause.
ArborDb: a Rust document store with constant-time field reads
ArborDb is a Rust document store using zero-copy, offset-table blobs to give constant-time field reads regardless of record size.
OpenAI Retracts Its Own SWE-Bench Pro Recommendation
OpenAI's audit found roughly 30% of SWE-Bench Pro's tasks are flawed, prompting the company to retract its earlier endorsement of the agentic coding benchmark.
Why AI Orchestration Beats Bigger Context Windows
Massive context windows didn't fix AI. With models scoring under 1% on ARC-AGI-3, winning teams now engineer the system around the model, not just the model.