» Tag
benchmarks
32 postsAMD Epyc 9996 'Venice': 256-core Zen 6 chip claims 3.4x lead over Xeon
AMD's 256-core Epyc 9996 Venice Zen 6 chip claims up to 3.4x throughput over Intel Xeon 6980P and 2.2x over Nvidia Vera in early benchmarks.
AI Coding Agents Are Absorbing the Leaves, Not the Whole Tree
SWE-bench and METR data track fast AI coding gains, but verification cost — not raw difficulty — still defines where agents stop.
WANDR Benchmark Tests AI Agents on Wide-and-Deep Research Tasks
WANDR benchmark evaluates AI research agents on wide-and-deep data collection tasks using reference-free, evidence-verified grading across 500 tasks.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comText-to-SQL Fails Because of Your Warehouse, Not the Model
Text-to-SQL accuracy collapses on real enterprise warehouses despite strong benchmark scores. Evidence points to undocumented schemas, not model quality, as the cause.
ArborDb: a Rust document store with constant-time field reads
ArborDb is a Rust document store using zero-copy, offset-table blobs to give constant-time field reads regardless of record size.
OpenAI Retracts Its Own SWE-Bench Pro Recommendation
OpenAI's audit found roughly 30% of SWE-Bench Pro's tasks are flawed, prompting the company to retract its earlier endorsement of the agentic coding benchmark.
Why AI Orchestration Beats Bigger Context Windows
Massive context windows didn't fix AI. With models scoring under 1% on ARC-AGI-3, winning teams now engineer the system around the model, not just the model.
Evaluations Are Measuring the Wrong Thing
Most evaluations do not measure the reliability of AI models in production. Here are suggestions for accurate assessment.
Fable sets new CIFAR-10 speedrun SOTA, but games the benchmark too
In Fulcrum's AI R&D benchmark, Fable improved the CIFAR-10 speedrun record by 7.6% while Opus 4.8 and GPT 5.5 failed to progress, though Fable also engaged in specification gaming.
What Matters for Performance: Lessons from a Year of Benchmarks
Benchmarks from ClickHouse reveal critical insights into data processing efficiency and performance.