» Tag
ai
105 posts7 Regression Tests Every AI Agent Should Pass Before Deploy
Discover essential regression tests for AI agents. Ensure a successful deployment with these key evaluations.
Sparse Policy Selection in RL for LLM Reasoning, Not Capability Learning
Reinforcement learning enhances LLM reasoning by focusing on sparse policy selection rather than teaching new capabilities.
AI Can Now Design Functional Viruses
Stanford researchers designed 16 AI-generated viruses capable of infecting antibiotic-resistant bacteria.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comStealing Reasoning Traces from Proprietary LLM APIs
We found a way to extract reasoning traces from frontier AI APIs, highlighting significant security risks and potential data leaks for engineers.
Open Model Enhanced for Scientific Inquiry
The open model developed by Loka and Arcee AI enhances scientific inquiry through tool use and logical reasoning.
NVIDIA Alpamayo 2 Super: 34B Open Vision-Language-Action Model
NVIDIA introduces Alpamayo 2 Super, a 34B open vision-language-action model for robotaxis.
I tested my own security tool and found four bugs
Agentmetry is a flight recorder for AI coding agents. In this entry, I discuss four bugs found and their implications for security tools.
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Mind viruses are self-propagating ideas in multi-agent systems. This study explores their risks and implications for AI agent design.
Cursor Launches Origin Code Hosting Platform Amid GitHub Outage
Cursor launched its Origin code hosting platform during a GitHub outage, aiming to transform source code management with AI integration.
Vero: Can AI Agents Build Formally Verified Software Repositories?
Vero benchmarks AI agents' capabilities in joint implementation and proof synthesis for software repositories.