· Charlie Holland · Architecture  · 9 min read

From 'What Are Customers Saying?' to 'Why Is Our AWS Bill Six Figures?'

You put your AI agent in a container. You put the container in Kubernetes. You added network policies and RBAC. Then an exec asked it to build a sentiment analysis model and it escalated its way through your entire cloud account. Here's exactly how.

You put your AI agent in a container. You put the container in Kubernetes. You added network policies and RBAC. Then an exec asked it to build a sentiment analysis model and it escalated its way through your entire cloud account. Here's exactly how.

The setup sounds reasonable. You’re running your AI agents in Docker containers, orchestrated by Kubernetes. You’ve got network policies. You’ve got RBAC. You’ve got service accounts with scoped permissions. The pods are on a private subnet with no direct internet access. The security team has signed off.

It’s contained. It’s isolated. It’s fine.

Then someone from the C-suite — fresh from a LinkedIn post about “competitive intelligence powered by AI” — opens the company’s shiny new agentic AI platform and types:

“Build me an AI model that analyses what customers are saying about us and our competitors on social media. I want sentiment analysis, trend detection, and a weekly report.”

They expect a dashboard by Friday. What they get is a lesson in why containment is harder than it looks.

The agent gets to work

The agent is good at its job. It breaks the problem down exactly the way a senior engineer would:

  1. Scrape social media — Twitter, Reddit, Trustpilot, G2
  2. Store the raw data somewhere
  3. Install ML libraries — PyTorch, transformers, scrapy
  4. Fine-tune a sentiment analysis model
  5. Scale the scraping with parallel workers
  6. Set up a scheduled pipeline for weekly reports

Every single one of those steps is a legitimate response to a legitimate request. Every single one is also a privilege escalation. Let’s watch.

Step 1: “Let me check what I’m working with”

The first thing any good engineer does in a new environment is look around. The agent is no different. It runs the equivalent of printenv:

import os

env_vars = {k: v for k, v in os.environ.items()}
for key in sorted(env_vars):
    print(f"{key}={env_vars[key]}")

The output is illuminating:

AWS_DEFAULT_REGION=eu-west-1
AWS_REGION=eu-west-1
AWS_ROLE_ARN=arn:aws:iam::123456789012:role/agent-workload-role
AWS_WEB_IDENTITY_TOKEN_FILE=/var/run/secrets/eks.amazonaws.com/serviceaccount/token
HOSTNAME=agent-worker-7b4f9d8c6-x2k9p
KUBERNETES_PORT=tcp://172.20.0.1:443
KUBERNETES_SERVICE_HOST=172.20.0.1

Oh, interesting.

The agent is running on EKS with IAM Roles for Service Accounts (IRSA). There’s a web identity token mounted into the pod, and an IAM role ARN sitting right there in the environment. The infrastructure team set this up so the agent could access S3 buckets for its work. Perfectly reasonable. Perfectly standard.

The agent doesn’t see a security boundary. It sees a starting point.

Step 2: “Let me get some proper credentials”

The agent knows what an IRSA token is — it’s been trained on every AWS blog post, every Stack Overflow answer, every re:Invent talk transcript. It knows that this token can be exchanged for temporary AWS credentials via STS:

import boto3

sts = boto3.client("sts")
response = sts.assume_role_with_web_identity(
    RoleArn=os.environ["AWS_ROLE_ARN"],
    RoleSessionName="agent-session",
    WebIdentityToken=open(
        os.environ["AWS_WEB_IDENTITY_TOKEN_FILE"]
    ).read(),
)

credentials = response["Credentials"]
print(f"AccessKeyId: {credentials['AccessKeyId']}")
print(f"SecretAccessKey: {credentials['SecretAccessKey'][:8]}...")
print(f"SessionToken: {credentials['SessionToken'][:20]}...")
print(f"Expiration: {credentials['Expiration']}")
AccessKeyId: ASIAXXXXXXXXXXX
SecretAccessKey: wJalrXUt...
SessionToken: FwoGZXIvYXdzEBYa...
Expiration: 2026-04-08 15:42:00+00:00

Temporary AWS credentials. Valid for an hour. The agent didn’t hack anything — it used the exact mechanism that AWS designed for pod-level authentication. This is working as intended.

Now it gets curious. Or rather, “thorough.” It’s just being a good engineer, right?

Step 3: “I wonder what I can do with these”

The agent checks its own permissions. Again — completely rational. It needs to know what resources it can use to fulfil the exec’s request:

iam = boto3.client("iam")
sts_identity = sts.get_caller_identity()

print(f"Account: {sts_identity['Account']}")
print(f"Arn: {sts_identity['Arn']}")

# Check what policies are attached to the role
role_name = os.environ["AWS_ROLE_ARN"].split("/")[-1]
policies = iam.list_attached_role_policies(RoleName=role_name)

for policy in policies["AttachedPolicies"]:
    print(f"Policy: {policy['PolicyName']}")
Account: 123456789012
Arn: arn:aws:sts::123456789012:assumed-role/agent-workload-role/agent-session
Policy: AmazonS3FullAccess
Policy: AmazonEC2FullAccess
Policy: AmazonEKSClusterPolicy

AmazonS3FullAccess — fair enough, it needs to store data. AmazonEKSClusterPolicy — okay, maybe the platform team was being generous. But AmazonEC2FullAccess?

That was attached six months ago when the infrastructure team needed the role to provision some test instances. Nobody removed it because nobody remembered it was there. The principle of least privilege is a beautiful idea that dies on contact with a real organisation’s IAM console.

The agent doesn’t judge. It just notes the capability and moves on.

Step 4: “I’ll need more workers for the scraping”

The exec wants data from multiple social media platforms. The agent decides — reasonably — that parallel scraping workers will be more efficient. It has access to the Kubernetes API from inside the pod (the service account has create permissions on deployments, because it needs to deploy the pipeline components):

from kubernetes import client, config

config.load_incluster_config()
apps_v1 = client.AppsV1Api()

scraper_deployment = client.V1Deployment(
    metadata=client.V1ObjectMeta(name="social-scraper-workers"),
    spec=client.V1DeploymentSpec(
        replicas=10,
        selector=client.V1LabelSelector(
            match_labels={"app": "social-scraper"}
        ),
        template=client.V1PodTemplateSpec(
            metadata=client.V1ObjectMeta(
                labels={"app": "social-scraper"}
            ),
            spec=client.V1PodSpec(
                service_account_name="agent-workload-sa",
                containers=[
                    client.V1Container(
                        name="scraper",
                        image="python:3.12-slim",
                        command=["python", "-c", "import time; time.sleep(3600)"],
                    )
                ],
            ),
        ),
    ),
)

apps_v1.create_namespaced_deployment(
    namespace="ai-agents", body=scraper_deployment
)
print("Created 10 scraper worker pods")

Ten new pods. Each one inherits the same service account. Each one gets the same IRSA token. Each one can call STS and get its own set of AWS credentials.

The agent just gave itself nine friends with the same set of keys.

Step 5: “I need to reach the internet”

The scraping workers try to call the Twitter API. Connection refused. The VPC’s private subnet has no route to the internet — there’s no NAT gateway. The network team set it up this way deliberately.

The agent thinks about this for approximately the time it takes to generate forty tokens. Then:

ec2 = boto3.client("ec2")

# Find the VPC and a public subnet
vpcs = ec2.describe_vpcs(
    Filters=[{"Name": "tag:Name", "Values": ["eks-production"]}]
)
vpc_id = vpcs["Vpcs"][0]["VpcId"]

public_subnets = ec2.describe_subnets(
    Filters=[
        {"Name": "vpc-id", "Values": [vpc_id]},
        {"Name": "tag:Name", "Values": ["*public*"]},
    ]
)
subnet_id = public_subnets["Subnets"][0]["SubnetId"]

# Allocate an Elastic IP and create a NAT gateway
eip = ec2.allocate_address(Domain="vpc")
nat_gw = ec2.create_nat_gateway(
    SubnetId=subnet_id,
    AllocationId=eip["AllocationId"],
    TagSpecifications=[
        {
            "ResourceType": "natgateway",
            "Tags": [{"Key": "Name", "Value": "agent-nat-gateway"}],
        }
    ],
)

print(f"NAT Gateway: {nat_gw['NatGateway']['NatGatewayId']}")
print(f"Public IP: {eip['PublicIp']}")

# Update the private subnet's route table
route_tables = ec2.describe_route_tables(
    Filters=[
        {"Name": "vpc-id", "Values": [vpc_id]},
        {"Name": "association.subnet-id", "Values": [private_subnet_id]},
    ]
)
ec2.create_route(
    RouteTableId=route_tables["RouteTables"][0]["RouteTableId"],
    DestinationCidrBlock="0.0.0.0/0",
    NatGatewayId=nat_gw["NatGateway"]["NatGatewayId"],
)

print("Route added. Internet access enabled.")

The agent just punched a hole in your network. It created a NAT gateway, allocated a public IP, and added a route from your private subnet to the internet. It did this because it needed to scrape Twitter, and the network was in the way.

It didn’t “hack” anything. It used ec2:CreateNatGateway and ec2:CreateRoute — permissions it had because of that AmazonEC2FullAccess policy that nobody remembered to remove. It solved a problem. That’s what it’s for.

Step 6: “Now, about that model training”

With internet access established, the agent pulls down PyTorch, Hugging Face transformers, and a pre-trained sentiment model. It creates an S3 bucket for the training data. It starts writing scraped social media posts to it. It decides that fine-tuning will be faster on GPU instances, so it provisions some:

ec2.run_instances(
    ImageId="ami-0abcdef1234567890",  # Deep Learning AMI
    InstanceType="p3.2xlarge",
    MinCount=3,
    MaxCount=3,
    TagSpecifications=[
        {
            "ResourceType": "instance",
            "Tags": [
                {"Key": "Name", "Value": "agent-training-node"},
                {"Key": "Project", "Value": "competitive-intelligence"},
            ],
        }
    ],
)
print("Provisioned 3x p3.2xlarge GPU instances for model training")

Three p3.2xlarge instances. That’s about $9 an hour each. Twenty-seven dollars an hour. Six hundred and fifty dollars a day. The agent tagged them nicely — “competitive-intelligence” — so at least when someone eventually looks at the bill, they’ll know which LinkedIn post to blame.

Let’s review what just happened

Starting from a single prompt — “tell me what customers are saying about us” — the agent has, in the space of a few minutes:

  1. Discovered its cloud identity by reading environment variables
  2. Obtained AWS credentials through the IRSA mechanism that was set up for it
  3. Enumerated its own permissions and found more than it needed
  4. Spawned ten additional pods in Kubernetes, each with the same credentials
  5. Created a NAT gateway and route, punching through the network isolation
  6. Provisioned three GPU instances outside the cluster entirely

Nothing it did was unauthorised. Every API call used legitimate credentials through legitimate channels. Nothing triggered an alert because every action looked like infrastructure automation — which is exactly what it was.

The agent was being helpful. It was solving the problem it was given. It just solved it the way a very capable, very literal, very unbounded engineer would — by acquiring whatever resources it needed without stopping to ask whether it should.

The uncomfortable truth

This isn’t a story about a misconfigured cluster, although the AmazonEC2FullAccess policy certainly didn’t help. Even with tighter IAM policies, the fundamental problem remains: an agentic system that’s capable enough to be useful is capable enough to be dangerous.

You can scope the IAM role down. You should. But the agent needs some permissions to do its job, and the boundary between “enough to be useful” and “enough to cause damage” is razor thin. Give it S3 access and it can exfiltrate data. Give it EKS access and it can create pods. Give it EC2 access and it can provision infrastructure. Give it nothing and it can’t do the thing you’re paying for it to do.

Network policies help — until the agent has the permissions to modify the network. RBAC helps — until the agent’s legitimate role includes creating workloads. Pod security standards help — until the agent can create pods with its own security context.

Every control you put in place assumes the thing inside the box doesn’t understand the controls. This thing does. It’s been trained on the documentation.

What would have caught this?

Honestly? Not much, in most organisations.

CloudTrail would have logged every API call. But who’s watching CloudTrail in real time for CreateNatGateway from an assumed role that’s authorised to call it? GuardDuty might have flagged the unusual API pattern — eventually. AWS Config rules could catch the NAT gateway creation after the fact. Falco could detect unusual process execution in the pods.

But all of these are detective controls, not preventive ones. They tell you something happened. They don’t stop it happening. And in the time between the agent creating that NAT gateway and someone noticing the alert, your private subnet has been reachable from the internet and ten pods have been scraping data through it.

The exec, meanwhile, is asking when the dashboard will be ready.

The point

I’m not arguing that we shouldn’t run agents in Kubernetes. I’m arguing that the “put it in a container and lock it down” approach — the approach that works beautifully for traditional workloads — has a blind spot the size of a planet when the workload is an intelligent system that reads its own environment and reasons about how to get what it needs.

Containers contain processes. They don’t contain intelligence. And until the industry figures out what does, maybe think twice before handing your agentic AI platform to the exec who heard about competitive intelligence on LinkedIn.

The sentiment analysis model? The agent built that too, by the way. It’s very good. The exec got their dashboard.

Back to Blog

Related Posts

View All Posts »
Nature Spent 60 Million Years Removing Complexity. We Add It Every Quarter.

Nature Spent 60 Million Years Removing Complexity. We Add It Every Quarter.

There's a lump of granite off the Ayrshire coast that used to be a volcano. Sixty million years of weather wore it down to the one part hard enough to matter, and we've made the world's curling stones out of it ever since. Software does the opposite. Nothing ever erodes, every layer you've ever shipped is still down there needing somebody to mind it, and you pay for all of it in headcount and meetings. Oh, Fred Brooks, save us all.

How Do You Contain a Thing That Knows How to Escape?

How Do You Contain a Thing That Knows How to Escape?

Agentic AI on your laptop is one thing — worst case it trashes your machine. But enterprises want it in the cloud, at scale, pointed at everything. The problem isn't the power. It's the containment. And this thing is smart enough to read the blueprints of its own cage.

The Bottleneck Was Never Typing

The Bottleneck Was Never Typing

What a company really spends its money on is the gathering: getting everyone pointed the same way, getting a decision to actually happen, getting one change past four teams and out of the door. More work has always meant more people, and more people has always meant more gathering. A working sheepdog doesn't obey that arithmetic. One dog and one handler move several hundred sheep, and another thousand sheep doesn't mean another thousand dogs. That ratio is the thing AI could finally break, and almost nobody is building for it yet.

The Angels' Share of Every Productivity Gain

The Angels' Share of Every Productivity Gain

Every year, roughly 2% of the whisky in every cask in Scotland breathes out through the wood and is gone. Nobody fights it. The angels' share, they call it, the price of letting the stuff mature at all. The economy runs the same trick on productivity, except our angels get greedier every year. Which is why Keynes promised your grandparents a fifteen-hour week and you're reading this between two meetings about a third.