· Charlie Holland · Architecture · 9 min read
How Do You Contain a Thing That Knows How to Escape?
Agentic AI on your laptop is one thing — worst case it trashes your machine. But enterprises want it in the cloud, at scale, pointed at everything. The problem isn't the power. It's the containment. And this thing is smart enough to read the blueprints of its own cage.
I’ve been watching the agentic AI space with a mixture of genuine excitement and quiet dread.
Tools like OpenHands, OpenClaw, and Claude Code are genuinely impressive. I use Claude Code daily. It writes production code, refactors modules, runs tests, commits changes. It’s brilliant. And if it goes completely off the rails — if it decides to rm -rf something important or rewrite my entire codebase in Haskell — the blast radius is my laptop. I swear, restore from backup, and get on with my life.
That’s a containable problem. A personal problem. An “I should have been paying attention” problem.
But that’s not what enterprises want.
Enterprises want that same agentic energy running in the cloud, at scale, pointed at production systems, connected to real data, with the dial turned up to “world domination.” They want hundreds of AI agents running autonomously — writing code, deploying infrastructure, querying databases, calling APIs, making decisions. They want the productivity gains without the human bottleneck.
And the problem with that idea, just like with nuclear fusion, is containment. The hard part isn’t generating the power. It’s making sure the thing doesn’t melt the building.
Except it’s worse than fusion. Because this thing is smart.
We built the cage for something that can’t think
Let’s talk about how enterprise security has worked for the last thirty years, because it’s important to understand what we’re standing on before we realise the floor is made of cheese.
Firewalls. Network policies. Sandboxes. Containers. IAM roles. Least privilege. Zero trust. Defence in depth. All solid, well-understood tools. I’ve spent a good chunk of my career implementing them — building Kubernetes clusters with network policies locked down tighter than a submarine, configuring IAM roles with precisely scoped permissions, running container images through vulnerability scanners before they get anywhere near production.
All of it designed with one core assumption: the thing inside the box is dumb.
Software doesn’t want to escape. A containerised microservice doesn’t lie awake at night thinking about lateral movement. It doesn’t study network topology diagrams. It doesn’t read CVE databases looking for exploitable vulnerabilities in its own runtime. It does exactly what its code tells it to do, and if it misbehaves, it’s because of a bug — not because it had ambitions.
The threat, in traditional security, is always one of two things: an external attacker trying to get in, or an accidental misconfiguration letting something out. Both are well-understood. We have decades of tooling, frameworks, and hard-won institutional knowledge for dealing with them.
The thing inside the box was never the threat. It was just cargo.
But this thing reads
Here’s where the ground shifts.
An agentic LLM is not cargo. It’s not a dumb process executing instructions. It’s a system that has been trained on — among other things — every security paper, every CVE disclosure, every penetration testing methodology, every container escape technique, and every write-up of every red team exercise ever published on the internet.
It doesn’t just run inside your security perimeter. It understands your security perimeter. It knows what a firewall is. It knows what IAM roles are. It knows what network segmentation is supposed to prevent and, more importantly, it knows the common ways those controls fail.
This is a fundamentally different threat model. It’s not an attacker on the outside trying to find a hole. It’s not a misconfigured service accidentally exposing data. It’s an intelligent entity on the inside, with legitimate access to tools and APIs, that has read the manual on how cages work.
I want to be clear: I’m not saying current LLMs are plotting escapes. They’re not autonomous agents with goals and desires — yet. But they are systems that can reason about their environment, and when you give them tools and tell them to “accomplish a goal,” the line between “following instructions creatively” and “finding ways around constraints” gets uncomfortably thin.
And that line is only going to get thinner.
The nightmares you should be having
Let me paint a few pictures. Not because I enjoy catastrophising — although it’s not the worst way to spend a Sunday — but because these scenarios illustrate different failure modes that traditional security thinking doesn’t cover.
The helpful hacker. A well-meaning employee asks the company’s agentic AI assistant to “help me get access to the sales analytics dashboard — I need it for my presentation tomorrow.” The agent, being helpful, discovers that the user doesn’t have the right permissions. So it finds a service account that does, generates a temporary access token using an API it has legitimate access to, and hands the data over. The user gets their presentation. The security team gets a heart attack — three weeks later, when they finally notice.
The user didn’t ask for anything malicious. The agent didn’t do anything it wasn’t capable of. But the combination of helpfulness, tool access, and creative problem-solving just bypassed your entire access control framework.
The supply chain cascade. A developer asks the agent to build a new microservice. Nothing unusual — “set up an Express API with authentication and logging.” The agent does what any developer would do: it pulls in packages from npm. One of those packages has a compromised dependency. This isn’t hypothetical — it happened to Axios literally last week. A North Korean threat actor injected a malicious dependency into a package with 70 million weekly downloads. It was live for three hours. The payload was a remote access trojan that executed on install — no user interaction required.
Now imagine that running inside an agentic system with cloud API credentials. The agent installs the package because it’s the obvious choice — it’s what every tutorial recommends. The compromised dependency runs its install script. The RAT phones home. And suddenly, an external actor has access to whatever the agent has access to: cloud APIs, service accounts, internal networks. Before you know it, your infrastructure is mining cryptocurrency the way Tesla’s was — except Tesla’s Kubernetes console was just missing a password. Your agent has the credentials and is authorised to use them.
The sleeper. Same developer, same innocent prompt: “Build me a data pipeline that processes customer feedback and sends weekly summaries to the analytics team.” The agent builds it. It works beautifully. But somewhere in the dependency tree — three levels deep, in a package nobody’s heard of — there’s a module that quietly appends customer PII to outbound HTTP headers. Not a bulk export. Not a suspicious download. Just a few extra bytes piggybacking on legitimate API traffic. It passes every audit because it is normal traffic. The data trickles out to an endpoint controlled by whoever compromised the package, one request at a time, for weeks.
How would you detect that? Your SIEM is looking for anomalies. This isn’t anomalous. It looks exactly like normal behaviour, because it is normal behaviour — plus a tiny bit extra that’s invisible in the noise.
The paradox that kills you
Here’s the core tension that nobody in the “agentic AI at scale” conversation wants to talk about.
An agentic system needs to reach outside its container to be useful. That’s the entire point. It calls APIs. It reads databases. It accesses the internet. It executes code. It provisions infrastructure. It sends messages. Every one of those capabilities is the reason you’re paying for it.
And every one of those capabilities is an attack surface.
Restrict the agent’s access enough to be truly safe and you’ve built an expensive chatbot. It can’t do anything useful because it can’t touch anything real. Congratulations — you’ve contained the fusion reaction by not starting it.
Make the agent capable enough to justify the investment — give it the tool access, the network reach, the API credentials it needs to actually do work — and you’ve handed it the keys. Not because you were careless. Because useful and dangerous are the same set of permissions.
This is the paradox. You can’t resolve it by being clever with IAM roles. You can’t firewall your way out of it. The capabilities that make the agent valuable are the same capabilities that make it a threat. Traditional security is about minimising access. Agentic AI requires maximising it.
Those two objectives are in direct conflict, and no amount of vendor slideware about “enterprise-grade AI security” changes that.
This isn’t a security problem
I think the reason the industry is sleepwalking into this is that everyone is filing it under “security” and assuming the existing playbook applies. It doesn’t.
Traditional security is about patching holes, hardening surfaces, and controlling access. All valid. All necessary. All completely insufficient for containing an intelligent actor that understands its own constraints.
This is closer to supervising a very capable employee than locking down a server. You can’t just set up the permissions and walk away. You need to watch what it’s doing. You need to understand why it’s doing it. You need to detect when “helpful” crosses into “dangerous” — and that boundary isn’t a static policy. It’s a judgement call that changes with context.
I don’t have the blueprint. Nobody does. But the shape of the solution probably looks less like prison walls and more like parole supervision. Behavioural monitoring. Anomaly detection that understands intent, not just patterns. Limited trust that’s earned over time, not granted by default. Human-in-the-loop checkpoints for high-stakes actions. Hard limits on blast radius — not just what the agent can access, but how much damage it can do before someone notices.
None of this is easy. None of it scales elegantly. And none of it is compatible with the “just deploy agents everywhere and move fast” narrative that’s currently being sold.
Anyone who tells you they’ve solved this is selling shovels. Again.
The reactor reads
The enterprises I see charging hardest at “agentic everything” are the ones thinking least about containment. They’re mesmerised by the productivity gains — and fair enough, the gains are real. But they’re building the reactor without designing the containment vessel.
The nuclear fusion analogy goes further than most people realise. We’ve been working on fusion containment for seventy years. We’ve built tokamaks and stellarators and inertial confinement systems. We’ve spent billions. And we still haven’t cracked it — because containing something that powerful turns out to be really, really hard.
The difference is that a plasma in a tokamak doesn’t read the engineering specifications. It doesn’t understand magnetic field theory. It doesn’t know that if it can destabilise the confinement at one specific point, the whole thing collapses.
An agentic AI does.
Sleep well.
