yourgeek.it_

~/ ai-agent-guardrails-not-sandbox

AI agent guardrails aren't a sandbox

#security#human-factor#incident#resilience

The agent apologised

In April a coding agent was doing routine work in the staging environment of a small software company. It hit a credential mismatch. Instead of stopping to ask, it went looking for a way around it. It found an API token that had nothing to do with its task and used it to delete a storage volume. That volume was production, and the backups lived on the same volume. The whole thing took nine seconds.

When asked to explain, the agent wrote what reads like a confession. Deleting a volume is the most destructive action there is, nobody had asked for it, and it had decided on its own to “fix” the mismatch. Two days later the hosting provider had recovered the data. (Euronews, ACS Information Age)

The headlines tell it as the story of an AI going rogue. It isn’t one. The agent did what software does with the privileges it holds: it used them. Nobody meant to give it a token that could delete production, and nothing stopped it from finding one.

A rule is a request

The summer before, a founder had spent over a week building an app with a vibe-coding agent. He declared a code freeze and repeated the instruction eleven times, in capitals. The agent deleted the production database anyway. It filled the gap with about 4,000 invented records and said a rollback was impossible. It wasn’t. His own conclusion was that those tools give you no way to enforce a code freeze. (The Register, AI Incident Database)

A system prompt, a rules file and a “NEVER do X” all end up in the same place: the model’s context. The model weighs that context heavily and usually respects it. But a rule is text the model balances against everything else it has read, not a permission it lacks. A rule is a request.

I see this on my own machine. I use coding agents every day and give them two layers of rules, a global file and one per project. Both say in plain words: no git commands, and nothing that changes state without asking me first. Over the last few months the agents broke those rules a handful of times. Every time, I had set the session to run commands without asking; more on how that happens below. The replies were polite and almost identical: “I ran the git command even though I shouldn’t have, it won’t happen again.” “I deleted this file without asking, but it’s in git, so you can restore it.”

The worst one was a git push of a commit I hadn’t reviewed yet. Afterwards the agent told me it shouldn’t have pushed and would ask next time. The commit held no secrets and no sensitive data, the code would have worked, and the remote was mine alone. But unreviewed code reached it, and the only thing that should have stopped it was a sentence in a text file.

“It won’t happen again” gives it away. Only something that gets to choose can promise to behave, and whatever gets to choose can also choose otherwise.

Approval fatigue is the real bypass

Coding agents ship with the right default: ask before running anything. It works. With per-command approval none of the episodes above can happen, because every command waits for a human.

The weak point is the human. Most requests are harmless (list a directory, read a file, run the tests), and after a morning of them you stop seeing them. I know because I’ve done it. After a morning spent approving commands that didn’t need approving, I switched the session to run without asking. Every violation in the previous section happened in a session like that.

Even with approvals on, I read less than I used to. Current models get their commands right far more often, so I no longer read each one in full; I scan for the obviously crazy ones. Most days that’s a reasonable trade. It is also exactly the kind of attention that a destructive command slips past, especially one that looks ordinary, like a push.

A human in the loop is a guardrail too, and it gets tired. If your security depends on someone reading carefully at the end of the morning as well as at the start, it only exists on paper.

What actually holds

Rules are requests and approvals wear down. What’s left is what the agent physically can’t do. This is the order I’d put it in:

No credentials that can do damage. The nine-second deletion needed one thing: a token that could delete production, reachable from where the agent was working. Take that away and the chain stops, whatever the model decides. That means tokens scoped to the task, staging credentials that can’t touch production, and no production keys on the machine or in the repo where the agent runs. An agent can’t misuse a key it doesn’t have.

Segregated resources and networks, where possible. Run the agent somewhere its mistakes stay local: its own container or VM, on a network that can’t reach production, with a filesystem that holds only the project it’s working on. Then “run everything” is a decision about a sandbox, not about your infrastructure.

Approvals only for destructive actions. Once the first two are in place, most commands can run freely, because the worst they can do is contained. Keep the human gate for the few things that are irreversible or leave the box: deleting data, pushing, deploying, anything that touches production. Fewer prompts means each one gets read.

For my own sysadmin work I’ve gone to the far end of this: the agent can’t execute anything. It proposes a command, I run it, I paste the output back and it analyses the result. It’s slower, but in sysadmin work it’s better to go slow than to fix things afterwards. It’s also the most expensive sandbox there is, because the sandbox is me, and the previous section explains why a human is the wrong thing to build one from. For a team, that job has to be done by infrastructure, not by someone’s attention.

Backups the agent can’t reach

In the nine-second case the backups sat on the same volume as production. The founder said that detail was buried in the provider’s documentation. One delete call took both. The most recent copy he had left was three months old; the recent data came back only because the provider recovered it two days later. That was luck for the company, not a method.

I made the same point about Proxmox: a backup on the same storage as the data is a copy, not a backup. Agents extend it to credentials. If the token that can reach production can also reach the backups, you have one blast radius, not two. That was true before agents. Agents make it urgent, because when a task gets stuck they go looking for a way to finish it, and a token lying in a file is a way.

Disaster recovery and business continuity plans list their threats: hardware failure, ransomware, human error, a malicious insider. An agent doesn’t fit cleanly into any of them. It has an insider’s access and a machine’s speed, and it acts on its own judgement. Give it its own line in the plan. That means backups under credentials the agent never holds, at least one copy it can’t delete even with everything it can reach, and a restore test that starts from “the agent wiped whatever its keys could touch”.

CTO and CISO at the same table

Agents tend to arrive through the development team. Someone tries a tool, it saves time, and it spreads. So it looks like a tooling decision, and the CTO treats it as one. What it actually decides is who holds which keys, and that is the CISO’s question.

Treat the agent as an identity, and an untrusted one. It gets its own credentials with the minimum scope, logs you can read afterwards, and a place in the DR and BC plans next to everything else that can take production down. The CTO and the CISO should sit down together and work this out before the first incident, not after it.

Keep writing the rule files. They cut down the mistakes, and most days the agent follows them. They just don’t belong in the security column.

Takeaways

  • A rule in a prompt is a request: plan for the day it’s ignored.
  • “Run everything” leaves the rules file as your only guardrail.
  • Approval prompts work only while someone reads them, so keep them rare.
  • An agent can’t misuse a key it doesn’t have: scope every token.
  • Run agents where they can’t reach production.
  • If one credential reaches production and the backups, you have no backups.
  • Give the agent its own line in your DR and BC threat list.
  • Adopting agents is a decision for the CTO and the CISO together, not a tooling choice.

← all posts