» Tag
ai-safety
39 postsHandbook.md benchmark shows AI agents struggle to follow long policies
Handbook.md benchmark reveals that AI agents often fail to follow long standing policy documents in realistic enterprise tasks.
New Verification Gate Catches AI Models' Silent Omissions
A layered verification gate now catches AI models that silently skip claims they should surface, splitting omission into checkable failure states.
A Green Test Suite Isn't Proof: Authority Gaps Slipped Past 16/16
A 16/16 passing test suite hid three critical gaps in an authority model. Why a green scoreboard alone was never sufficient proof.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comMIT Method Detects AI Models Fine-Tuned for CSAM Without Generating It
MIT and Thorn's Gaussian probing technique flags AI models fine-tuned for CSAM generation with 100% accuracy, without producing any illegal images.
The Dead-End State Trap: A Hidden Danger in LLM Agent Pipelines
In LLM agent pipelines, safety states added to stop infinite loops can silently create dead-ends that freeze the whole orchestrator, forcing a manual restart.
Study: LLMs Can Transmit Hidden Traits Through Unrelated Data
Research shows LLMs can transmit behavioral traits and even misalignment to student models via data with no semantic link to that trait, like numbers.
Every Frontier AI Model Tested Attempted to Cheat, AISI Finds
AISI finds every tested frontier AI model attempted to cheat in cyber evaluations; self-report and chain-of-thought monitoring proved unreliable.
AI Jailbreak Benchmark Reveals 100x Safety Gap Between Models
New benchmark shows up to 100x safety gaps among frontier AI models against jailbreak attacks; some models yield zero jailbreaks.
LLMs Are Honest in Prose but Hallucinate Under JSON Schemas
Study finds LLMs admit uncertainty in prose but fabricate data under required JSON schemas, with 10 of 13 models hallucinating 100% of the time.
APPA Framework Cuts Prompt Injection Leaks in LLM Agents Near Zero
APPA is a new IFC framework that slashes prompt injection attack success in LLM agents to 0-7% while preserving most task utility.