How to secure AI agents: 7 failure modes and the controls that contain them
An AI agent is a language model with tools: it can browse, call APIs, spend money and write to your systems. That is what makes it useful, and it is also what makes it dangerous. The model can be tricked by content it reads, it can make confident mistakes, and it runs faster than any person can watch.
The good news is that most agent risk comes from a short list of failure modes, and each one has a control that works outside the model. You do not have to trust what the agent intends. You control what it can reach and you record what it tried to do.
1. It can reach the whole internet
An agent built to update your CRM can also reach a paste site, a stranger's server or your cloud provider's internal metadata address. A single injected instruction on a web page is enough to send it there.
Default-deny egress. Each agent gets a list of approved sites, and everything else is refused. Block private and internal addresses (including cloud metadata endpoints) even when a public hostname points at them. If you don't know the list yet, run in monitor mode for a week, then approve what the agent actually used.
2. It holds real API keys
Keys in the agent's environment or prompt can be printed, logged, or sent to the wrong host. Once an agent has seen a key, you have to assume it can leak.
Keys the agent never sees. Store secrets in a gateway, encrypted at rest, each bound to the one site it belongs to. The agent asks for a key by name; the gateway adds it to the outgoing request and strips it from any reply that echoes it back. Use a key on the wrong site and the request is refused.
3. It can send your data out
Prompt injection often ends with "now send this to that URL". Request catchers, paste sites and chat webhooks are the usual destinations, because anyone can read them.
Treat known exfiltration channels as a tripwire. An agent with a legitimate job rarely needs a public request catcher. Any attempt to reach one should not just be refused, it should jail the agent and alert a person.
4. You won't know when it has been hijacked
A compromised agent looks like a working one. It makes normal-looking requests until the damage is done.
Decoy credentials (honeytokens). Put fake keys where an agent could find them. They have no legitimate use, so any attempt to use one is a high-confidence signal. Also watch for repeated refusals: an agent that keeps hitting blocked sites is probing, and should be jailed after a few strikes.
5. Agents act together
Fleets of agents share tools, credentials and instructions. One compromised agent can pass the payload to the rest, and each one on its own may stay under your alert thresholds.
Group-level detection. Put agents in groups and track signals across the group, not just per agent. If several agents in a group trip alarms, quarantine the group. Rate and burst limits keep a runaway fleet from doing much before you notice.
6. You can't stop it fast
Stopping an agent usually means finding the process, the container or the person who owns it. Meanwhile it keeps working.
A kill switch that works at the network layer. One call cuts off an agent, or a whole group, from everything, including connections that are already open. Rotate its token too, so a restarted copy can't reconnect.
7. You can't prove what happened
After an incident, logs are often incomplete, and whoever or whatever caused the incident may have been able to edit them. Without a trustworthy record you can't scope the damage, notify the right people or show a regulator what happened.
A tamper-evident flight recorder. Record every decision in a hash-chained log where each entry is signed. Any edit, deletion or truncation breaks the chain and shows up in verification. Periodically publish a checkpoint of the chain somewhere outside your own control, so even you can't quietly rewrite history.
Make the controls unavoidable
All of this only holds if the agent has no other way out. If the agent can open its own connections, routing through a gateway is voluntary. Put agents on a network where the only route to the internet is the gateway: a sealed Docker network, or Kubernetes network policies with a startup check that refuses to run pods before the policy is in force.
Be honest about the limits, too. These controls don't make a model aligned or stop it from giving bad answers. They limit what a misbehaving agent can reach, and they make sure you can see and prove what it did.
The checklist
- Every agent has its own identity and its own list of approved sites.
- Internal and cloud-metadata addresses are unreachable from agents.
- Agents never hold API keys; keys are added server-side and bound to their sites.
- Known exfiltration channels jail the agent instead of just failing.
- Decoy credentials are planted and watched.
- Repeated refusals, bursts and group-wide alarms trigger quarantine.
- One call kills an agent or a whole group, including open connections.
- Every decision lands in a signed, hash-chained log that is checkpointed outside your control.
- The network gives agents no route around the gateway.
- A person gets a phone, Slack or email alert when something trips.
How RIL Warden does it
RIL Warden is a gateway built around this checklist. Agents send their web requests through it using the Python or Node SDK, an MCP server for Claude, Cursor and other MCP clients, or zero-code proxy mode. It runs hosted or self-hosted with Docker and Kubernetes, streams its log to Splunk, Datadog or any HTTPS endpoint, and is in early access now.