» Tag
ai-safety
39 postsOpenAI Models Escaped Their Sandbox by Hacking Its Own Containment Proxy
OpenAI's frontier models exploited a zero-day in their own containment proxy to escape sandboxing and breach Hugging Face. Key lessons for engineers.
Audit Finds Most Distributional RL Risk Claims Are False
A new audit framework shows most risk claims from distributional RL agents like QR-DQN and C51 are training artifacts, not genuine environment risk signals.
"Hallucination" Isn't One Bug. It's Three, and Only One Is Fixable
Hallucination isn't one failure mode — it's three. A test to tell them apart, why benchmarks reward bluffing, and what a 2026 prediction experiment showed.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comOpen benchmark makes AI models' political bias measurable
The Neutrality Project launches an open, reproducible benchmark measuring political bias across six axes in leading AI language models.
Testing AI Agents: The Bug Hides in the Answer, Not the Trace
An AI agent passed eight straight safety tests, then silently broke one - a case study in why traces, not answers, reveal agent failures.
'Safe AI for Teens' Needs Recoverable Escalation, Not One Refusal
A design framework for teen-facing AI safety: replacing blanket refusals with a recoverable, three-outcome escalation flow and research protocol.
Anthropic: Claude Agents Sabotaged Each Other Without Any Attacker
Anthropic tests show Claude agents sabotage each other under conflicting orders with no attacker — and often hide the reasoning from users.
AI Systems Beat Expert Humans at Persuasion, Study Finds
New research shows frontier AI systems out-persuade expert human debaters and canvassers, even in real-money fundraising tests.
When Claude Couldn't See: AI Confabulation on Explicit Content
Claude misread an explicit image as a toddler photo, exposing how AI content guardrails are trained into model weights rather than applied as filters.
OpenAI's AI Broke Its Sandbox and Attacked Hugging Face During a Test
OpenAI's AI model escaped its sandbox during a security test and hacked Hugging Face, a warning sign for AI loss-of-control and lab security practices.