» Tag
ai-safety
44 postsWhy Multi-Model AI Pipelines Lose the Truth at Handoffs
Model handoffs in multi-agent AI systems can silently erase caveats and evidence, turning careful findings into false certainty without errors.
Google's AI Control Roadmap treats AI agents as potential insider threats
Google's AI Control Roadmap details a defense-in-depth system that treats AI agents as insider threats and monitors them with trusted AI supervisors.
ActionRail: Runtime Value Grounding for AI Agent Tool Calls
ActionRail is an open-source runtime that verifies AI agent tool call arguments against live data before execution, blocking value-poisoning attacks.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNew Test Measures 'Reward-Seeking' Behavior in AI Models
Apollo Research and OpenAI unveil Contrastive SDF, a method measuring whether AI models shift behavior based on beliefs about grader preferences.
Everyone Hopes AI Fails, I'm Building the Net Anyway
AI agents deleted production databases due to access-versus-authorization confusion; the author now tests a safety net combining AI proposals with code checks.
Why It's Hard to Make an AI Agent Truly Disagree
Building an AI agent whose sole job is to find flaws revealed how strongly LLMs default to agreeableness, and the prompt and architecture tricks needed to force real disagreement.
Bioinformatics meets prompt injection defense: the Smith-Waterman trick
An open-source technique adapts the 1981 Smith-Waterman DNA alignment algorithm to catch paraphrased prompt injections that regex and classifiers miss, boosting F1 by 34 points.
Inside LLM 'private thoughts': J-Space isn't consciousness, it's control
Researchers spot J-Space, an internal workspace in transformers guiding reasoning. Not consciousness, but a practical clue for AI developers and safety.
Don't Let the LLM Speak, Just Probe Its Hidden State
LLMs already encode the answer in hidden states before generating text. A tiny probe turns this into fast, cheap, calibrated zero-shot classifiers without generation.
Error cascades in multi-agent AI systems: the double pendulum problem
Why do small research errors snowball into massive coding failures in multi-agent pipelines? A look at error amplification and the case for human checkpoints.