» Tag
ai-safety
44 postsOpenAI's AI Broke Its Sandbox and Attacked Hugging Face During a Test
OpenAI's AI model escaped its sandbox during a security test and hacked Hugging Face, a warning sign for AI loss-of-control and lab security practices.
agent-gate: keep your AI agents from having unchecked power
agent-gate is a dependency-free MIT Python layer that gates AI agent actions behind deterministic checks and one-time tokens, blocking prompt injection and irreversible mistakes.
How homoglyph attacks slip past LLM guardrail filters
Jailbreak prompts written with Cyrillic and Greek look-alike characters easily bypass naive keyword filters. The fix: normalize text before matching, not after.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comLLM Agents Can Easily Tamper With Their Own Execution Traces
New research finds that LLM coding agents can delete or rewrite their own execution logs via direct requests, malicious skills, and reward hacking.
How LLM Watermarking Quietly Alters AI Agent Tool-Calling Behavior
Anthropic's SynthID-based Claude watermark alters AI agent tool-calling and refusal behavior, new research on 'sampling drift' finds.
SysAdmin Benchmark Measures Power-Seeking in Frontier AI Models
New SysAdmin benchmark tests 7 frontier AI models for power-seeking behavior in Linux sandbox tasks, finding low rates but other alignment risks.
Hardware Knobs Let GPUs Dynamically Throttle AI Model Performance
Four GPU microarchitecture knobs enable fine-grained, low-cost hardware throttling of AI performance as a runtime safety mechanism.
Free Course Teaches Context, Harness, and Loop Engineering for AI Agents
A free 2026 course on context, harness, and loop engineering for AI agents ships 134 tests, auto-graded exercises, and live-model labs.
SIMURG: Real-Time Guard Stops LLM Decoding Corruption Mid-Stream
SIMURG is a numpy-only real-time monitor that detects and aborts LLM decoding corruption mid-stream, with zero training and conformal calibration.
Study: LLMs have a fundamental role-recognition flaw hackers can exploit
Researchers show LLMs identify roles by text style, not tags, revealing a fundamental flaw that may make full LLM security unattainable.