Prompt injection, including indirect injection from a web page or document an agent reads, can make an agent follow an attacker's instructions: leak secrets, call the wrong API, or post your data to an attacker's server. Input filtering helps but never catches everything. Warden takes the other half of the problem: containment.
Containment, not just detection
It cannot reach the attacker
Default-deny egress: a hijacked agent still can only connect to sites you approved. The attacker's drop site is simply blocked.
It cannot leak the key
Secrets are injected server-side only to their real host. An injected instruction to 'print your API key' returns nothing useful.
It cannot exfiltrate
Known exfiltration and C2 channels are refused, and oversized or suspicious egress is caught and can auto-jail the agent.
You see it happen
Every attempt is written to a tamper-proof, signed log, so you have a record of exactly what the injection tried to do.
Defense in depth
Use your model-side input filters AND Warden's action-side containment together. The filters reduce how often an agent is tricked; Warden makes sure that when one slips through, the agent cannot actually do harm. Learn the full set of failure modes in the guide to securing AI agents.
Frequently asked questions
- Can you fully prevent prompt injection?
- No tool can guarantee a model is never tricked, especially with indirect injection from content the agent reads. The reliable defense is containment: ensure a tricked agent still cannot reach, leak, or exfiltrate. That is what RIL Warden enforces.
- What is indirect prompt injection?
- It is when malicious instructions are hidden in content an agent consumes - a web page, a PDF, an email - rather than typed by the user. The agent reads them as instructions. Warden assumes this will happen and limits what the agent can do as a result.
- How does Warden help if the model is already tricked?
- It enforces policy on the agent's actions, not the text. A hijacked agent is still blocked from unapproved destinations, never sees raw secrets, and is caught and jailed if it tries to exfiltrate - all recorded in a signed log.
- Is this a replacement for input guardrails?
- No, it is the complement. Guardrails reduce successful injections; Warden contains the ones that get through. Teams use both.
Related
Put a wall around your agents in 5 minutes
Free tier: 3 agents, 25,000 gateway requests a month, full signed history.
Start freeRead the guide