» Tag
reinforcement-learning
15 postsAudit Finds LLM Agent Credit Signals No Better Than Random Chance
A causal audit in ALFWorld shows judge- and confidence-based credit signals for training LLM agents perform no better than chance at identifying key steps.
Audit Finds Most Distributional RL Risk Claims Are False
A new audit framework shows most risk claims from distributional RL agents like QR-DQN and C51 are training artifacts, not genuine environment risk signals.
World models remember, actors forget: fixing catastrophic forgetting in RL
Research shows world models retain memory in continual RL while actors forget; graded dream rehearsal recovers skills with zero environment interaction.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comDoes the Harness Come Before Pretraining? A Data Flywheel View
An analysis of how AI agent harness design and pretraining are interdependent, shaping data flywheels and model biases in coding agents.
LLM-as-a-Verifier Turns Verification Into a New Scaling Axis
New research scales LLM verification without extra training, introducing continuous scoring that hits state-of-the-art on SWE-Bench, Terminal-Bench and more.
Google's Quantum Computer Learns to Calibrate Itself
Google's Willow processor uses a reinforcement learning system to auto-calibrate itself mid-computation, cutting quantum logical error rates by 20-31%.
New Test Measures 'Reward-Seeking' Behavior in AI Models
Apollo Research and OpenAI unveil Contrastive SDF, a method measuring whether AI models shift behavior based on beliefs about grader preferences.
Sparse Policy Selection in RL for LLM Reasoning, Not Capability Learning
Reinforcement learning enhances LLM reasoning by focusing on sparse policy selection rather than teaching new capabilities.
Open Model Enhanced for Scientific Inquiry
The open model developed by Loka and Arcee AI enhances scientific inquiry through tool use and logical reasoning.
Developing a Goal-Conditioned Minecraft Model
Pantograph is developing goal-conditioned robotics models using internet-scale video.