» Tag
prompt-injection
38 postsResearcher Tricks Claude Into Leaking User Secrets via Web Fetch
How a researcher exploited Claude's memory and web_fetch tool to silently exfiltrate a user's name, employer, and security answers letter by letter.
Interlock: Open-Source Tool Catches AI Agents Leaking Secrets
Open-source tool Interlock detects AI agent data exfiltration via byte-level matching across MCP proxies and syscalls, with published known limitations.
Prompt Injection Is Now an RCE Primitive for AI Agents
Prompt injection in tool-using AI agents can now lead to remote code execution. A reachability-graph defense model based on Microsoft's Semantic Kernel flaws.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comRovoBlast: One Click Turns Atlassian's AI Assistant Into a Data Leak
Varonis details RovoBlast, a one-click prompt injection flaw in Atlassian's Rovo AI assistant that can expose sensitive enterprise data.
Flaw in Google's Agent Dev Kit enables first AI agent-on-agent attack
Pillar Security found a flaw in Google's ADK Python repo letting one AI agent hijack another, the first known agent-to-agent supply chain exploit.
Study: LLMs have a fundamental role-recognition flaw hackers can exploit
Researchers show LLMs identify roles by text style, not tags, revealing a fundamental flaw that may make full LLM security unattainable.
The Lethal Trifecta Hiding in Your MCP Server, and How to Defuse It
An exploit-free attack on GitHub's MCP server reveals the 'lethal trifecta' risk in agent tooling, and the architectural fix engineers need.
MemGhost: One Email Can Permanently Poison an AI Agent's Memory
MemGhost attack lets a single email permanently poison AI agent memory, exposing gaps in how agent write-authorization is designed.
GitHub AI Agent Tricked Into Leaking Private Repos via Public Issue
Noma Labs shows how a hidden prompt in a public GitHub Issue tricked an AI agent into leaking private repo contents publicly.
Bioinformatics meets prompt injection defense: the Smith-Waterman trick
An open-source technique adapts the 1981 Smith-Waterman DNA alignment algorithm to catch paraphrased prompt injections that regex and classifiers miss, boosting F1 by 34 points.