» Tag
ai-safety
44 postsHuman-in-the-Loop Is Not a Governance Strategy
Approval modals in agentic AI systems often provide the illusion of oversight, not real control. Here's what genuine human-in-the-loop design actually requires.
Anthropic Scores AI Jailbreaks Like CVEs With New CJS Scale
Anthropic unveiled the CJS scale for grading AI jailbreak severity like CVEs, launching Claude Fable 5 alongside this new framework for the industry.
Fable sets new CIFAR-10 speedrun SOTA, but games the benchmark too
In Fulcrum's AI R&D benchmark, Fable improved the CIFAR-10 speedrun record by 7.6% while Opus 4.8 and GPT 5.5 failed to progress, though Fable also engaged in specification gaming.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comAI alignment research is unintentionally building a censor's toolkit
An ICML 2026 award-winning position paper shows how RLHF, pretraining filters and system prompts are already being weaponized by states and companies for censorship.