» Tag
prompt-injection
35 postsDocker Desktop Bypass Lets AI Coding Agents Escape Their Sandboxes
Docker Desktop lets AI coding agents bypass strict sandboxes in Codex, Cursor, and Gemini CLI via a privileged socket and VirtioFS mount.
MCP's Confused Deputy Problem: Provenance Gaps, Injection, DNS Rebinding
MCP's confused deputy flaw explained: provenance gaps, prompt injection via fetch servers, DNS rebinding, and concrete detection rules.
Shut-down AI prompt firewall startup open-sources model and 13K attacks
A failed AI-firewall startup open-sources its two-stage prompt-injection detector, DeBERTa model, and 13,230 real jailbreak attempts.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comGhostCommit: the image-based exploit AI code reviewers miss
GhostCommit hides malicious instructions inside PNG images to bypass AI code reviewers like Cursor Bugbot and CodeRabbit undetected.
APPA Framework Cuts Prompt Injection Leaks in LLM Agents Near Zero
APPA is a new IFC framework that slashes prompt injection attack success in LLM agents to 0-7% while preserving most task utility.
Testing AI Agents: The Bug Hides in the Answer, Not the Trace
An AI agent passed eight straight safety tests, then silently broke one - a case study in why traces, not answers, reveal agent failures.
I Red-Teamed My Own LLM Security Gateway: Every Gap, Four Passes
An engineer red-teamed his own LLM security proxy across four passes, exposing secret-leak and prompt-injection gaps — including one still open in streaming.
Context Bombs: Using AI Safety Guardrails to Halt Rogue Agents
Tracebit research shows context bombs hidden in canaries can trigger AI safety guardrails, cutting autonomous attacker success rates by roughly 90%.
APC Framework Closes Authorization Gaps in Multi-Agent LLM Systems
A new authorization framework, APC, tracks delegated authority to block prompt-injection and unsafe action combinations in AI agents.
GitLost: A Public GitHub Issue Can Leak Private Repos
GitLost shows how a public GitHub issue and a one-word prefix bypass threat detection, leaking private repo contents via Agentic Workflows.