» Tag
model-evaluation
3 postsAudit Finds Most Distributional RL Risk Claims Are False
A new audit framework shows most risk claims from distributional RL agents like QR-DQN and C51 are training artifacts, not genuine environment risk signals.
Mechanistic View Reveals How Bias Lives Inside LLM Judges
Study shows LLM-as-judge bias is encoded in activation geometry, enabling causal steering and better failure prediction than text-based methods.
A Production Checklist for Rolling Out Open-Weight AI Models
A practical rollout guide for teams adopting open-weight AI models, covering task contracts, eval sets, routing layers, and output validation.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com