MikeTrendsTrends right now

⬢github Python · 19 ★ +1 since we first saw it · pushed 15 h ago · MIT

rudratoshs/buried-injections

🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark.

buried-injections is a reproducible Python benchmark that runs 10 open-source prompt-injection detectors against 629 realistic AgentDojo attacks embedded in tool output, the way agent firewalls actually see them. It shows none catches most attacks without blocking safe traffic, and that tuning thresholds dramatically reshuffles rankings — Prompt Guard 2 goes from worst to best at a 2% false-alarm budget.

Why now: It was discussed on Hacker News because its results challenge widely used prompt-injection defenses like Meta's Prompt Guard 2, showing they catch as little as 1% of realistic attacks when the default thresholds are used.

Who it is for: Developers building AI agent firewalls, guardrails, or MCP/tool pipelines who need to choose and tune a prompt-injection detector.

llm-securityprompt-injectionbenchmarkai-agentspython

agentdojoai-agentsbenchmarkllm-securitymcpprompt-guardprompt-injection

Open on GitHub →

Stars over our 39 snapshots: 18 to 19, since 7 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/rudratoshs/buried-injections