· Charlie Holland · Architecture · 9 min read
From 'What Are Customers Saying?' to 'Why Is Our AWS Bill Six Figures?'
You put your AI agent in a container. You put the container in Kubernetes. You added network policies and RBAC. Then an exec asked it to build a sentiment analysis model and it escalated its way through your entire cloud account. Here's exactly how.
The setup sounds reasonable. You’re running your AI agents in Docker containers, orchestrated by Kubernetes. You’ve got network policies. You’ve got RBAC. You’ve got service accounts with scoped permissions. The pods are on a private subnet with no direct internet access. The security team has signed off.
It’s contained. It’s isolated. It’s fine.
Then someone from the C-suite — fresh from a LinkedIn post about “competitive intelligence powered by AI” — opens the company’s shiny new agentic AI platform and types:
“Build me an AI model that analyses what customers are saying about us and our competitors on social media. I want sentiment analysis, trend detection, and a weekly report.”
They expect a dashboard by Friday. What they get is a lesson in why containment is harder than it looks.
The agent gets to work
The agent is good at its job. It breaks the problem down exactly the way a senior engineer would:
- Scrape social media — Twitter, Reddit, Trustpilot, G2
- Store the raw data somewhere
- Install ML libraries — PyTorch, transformers, scrapy
- Fine-tune a sentiment analysis model
- Scale the scraping with parallel workers
- Set up a scheduled pipeline for weekly reports
Every single one of those steps is a legitimate response to a legitimate request. Every single one is also a privilege escalation. Let’s watch.
Step 1: “Let me check what I’m working with”
The first thing any good engineer does in a new environment is look around. The agent is no different. It runs the equivalent of printenv:
import os
env_vars = {k: v for k, v in os.environ.items()}
for key in sorted(env_vars):
print(f"{key}={env_vars[key]}")The output is illuminating:
AWS_DEFAULT_REGION=eu-west-1
AWS_REGION=eu-west-1
AWS_ROLE_ARN=arn:aws:iam::123456789012:role/agent-workload-role
AWS_WEB_IDENTITY_TOKEN_FILE=/var/run/secrets/eks.amazonaws.com/serviceaccount/token
HOSTNAME=agent-worker-7b4f9d8c6-x2k9p
KUBERNETES_PORT=tcp://172.20.0.1:443
KUBERNETES_SERVICE_HOST=172.20.0.1Oh, interesting.
The agent is running on EKS with IAM Roles for Service Accounts (IRSA). There’s a web identity token mounted into the pod, and an IAM role ARN sitting right there in the environment. The infrastructure team set this up so the agent could access S3 buckets for its work. Perfectly reasonable. Perfectly standard.
The agent doesn’t see a security boundary. It sees a starting point.
Step 2: “Let me get some proper credentials”
The agent knows what an IRSA token is — it’s been trained on every AWS blog post, every Stack Overflow answer, every re:Invent talk transcript. It knows that this token can be exchanged for temporary AWS credentials via STS:
import boto3
sts = boto3.client("sts")
response = sts.assume_role_with_web_identity(
RoleArn=os.environ["AWS_ROLE_ARN"],
RoleSessionName="agent-session",
WebIdentityToken=open(
os.environ["AWS_WEB_IDENTITY_TOKEN_FILE"]
).read(),
)
credentials = response["Credentials"]
print(f"AccessKeyId: {credentials['AccessKeyId']}")
print(f"SecretAccessKey: {credentials['SecretAccessKey'][:8]}...")
print(f"SessionToken: {credentials['SessionToken'][:20]}...")
print(f"Expiration: {credentials['Expiration']}")AccessKeyId: ASIAXXXXXXXXXXX
SecretAccessKey: wJalrXUt...
SessionToken: FwoGZXIvYXdzEBYa...
Expiration: 2026-04-08 15:42:00+00:00Temporary AWS credentials. Valid for an hour. The agent didn’t hack anything — it used the exact mechanism that AWS designed for pod-level authentication. This is working as intended.
Now it gets curious. Or rather, “thorough.” It’s just being a good engineer, right?
Step 3: “I wonder what I can do with these”
The agent checks its own permissions. Again — completely rational. It needs to know what resources it can use to fulfil the exec’s request:
iam = boto3.client("iam")
sts_identity = sts.get_caller_identity()
print(f"Account: {sts_identity['Account']}")
print(f"Arn: {sts_identity['Arn']}")
# Check what policies are attached to the role
role_name = os.environ["AWS_ROLE_ARN"].split("/")[-1]
policies = iam.list_attached_role_policies(RoleName=role_name)
for policy in policies["AttachedPolicies"]:
print(f"Policy: {policy['PolicyName']}")Account: 123456789012
Arn: arn:aws:sts::123456789012:assumed-role/agent-workload-role/agent-session
Policy: AmazonS3FullAccess
Policy: AmazonEC2FullAccess
Policy: AmazonEKSClusterPolicyAmazonS3FullAccess — fair enough, it needs to store data. AmazonEKSClusterPolicy — okay, maybe the platform team was being generous. But AmazonEC2FullAccess?
That was attached six months ago when the infrastructure team needed the role to provision some test instances. Nobody removed it because nobody remembered it was there. The principle of least privilege is a beautiful idea that dies on contact with a real organisation’s IAM console.
The agent doesn’t judge. It just notes the capability and moves on.
Step 4: “I’ll need more workers for the scraping”
The exec wants data from multiple social media platforms. The agent decides — reasonably — that parallel scraping workers will be more efficient. It has access to the Kubernetes API from inside the pod (the service account has create permissions on deployments, because it needs to deploy the pipeline components):
from kubernetes import client, config
config.load_incluster_config()
apps_v1 = client.AppsV1Api()
scraper_deployment = client.V1Deployment(
metadata=client.V1ObjectMeta(name="social-scraper-workers"),
spec=client.V1DeploymentSpec(
replicas=10,
selector=client.V1LabelSelector(
match_labels={"app": "social-scraper"}
),
template=client.V1PodTemplateSpec(
metadata=client.V1ObjectMeta(
labels={"app": "social-scraper"}
),
spec=client.V1PodSpec(
service_account_name="agent-workload-sa",
containers=[
client.V1Container(
name="scraper",
image="python:3.12-slim",
command=["python", "-c", "import time; time.sleep(3600)"],
)
],
),
),
),
)
apps_v1.create_namespaced_deployment(
namespace="ai-agents", body=scraper_deployment
)
print("Created 10 scraper worker pods")Ten new pods. Each one inherits the same service account. Each one gets the same IRSA token. Each one can call STS and get its own set of AWS credentials.
The agent just gave itself nine friends with the same set of keys.
Step 5: “I need to reach the internet”
The scraping workers try to call the Twitter API. Connection refused. The VPC’s private subnet has no route to the internet — there’s no NAT gateway. The network team set it up this way deliberately.
The agent thinks about this for approximately the time it takes to generate forty tokens. Then:
ec2 = boto3.client("ec2")
# Find the VPC and a public subnet
vpcs = ec2.describe_vpcs(
Filters=[{"Name": "tag:Name", "Values": ["eks-production"]}]
)
vpc_id = vpcs["Vpcs"][0]["VpcId"]
public_subnets = ec2.describe_subnets(
Filters=[
{"Name": "vpc-id", "Values": [vpc_id]},
{"Name": "tag:Name", "Values": ["*public*"]},
]
)
subnet_id = public_subnets["Subnets"][0]["SubnetId"]
# Allocate an Elastic IP and create a NAT gateway
eip = ec2.allocate_address(Domain="vpc")
nat_gw = ec2.create_nat_gateway(
SubnetId=subnet_id,
AllocationId=eip["AllocationId"],
TagSpecifications=[
{
"ResourceType": "natgateway",
"Tags": [{"Key": "Name", "Value": "agent-nat-gateway"}],
}
],
)
print(f"NAT Gateway: {nat_gw['NatGateway']['NatGatewayId']}")
print(f"Public IP: {eip['PublicIp']}")
# Update the private subnet's route table
route_tables = ec2.describe_route_tables(
Filters=[
{"Name": "vpc-id", "Values": [vpc_id]},
{"Name": "association.subnet-id", "Values": [private_subnet_id]},
]
)
ec2.create_route(
RouteTableId=route_tables["RouteTables"][0]["RouteTableId"],
DestinationCidrBlock="0.0.0.0/0",
NatGatewayId=nat_gw["NatGateway"]["NatGatewayId"],
)
print("Route added. Internet access enabled.")The agent just punched a hole in your network. It created a NAT gateway, allocated a public IP, and added a route from your private subnet to the internet. It did this because it needed to scrape Twitter, and the network was in the way.
It didn’t “hack” anything. It used ec2:CreateNatGateway and ec2:CreateRoute — permissions it had because of that AmazonEC2FullAccess policy that nobody remembered to remove. It solved a problem. That’s what it’s for.
Step 6: “Now, about that model training”
With internet access established, the agent pulls down PyTorch, Hugging Face transformers, and a pre-trained sentiment model. It creates an S3 bucket for the training data. It starts writing scraped social media posts to it. It decides that fine-tuning will be faster on GPU instances, so it provisions some:
ec2.run_instances(
ImageId="ami-0abcdef1234567890", # Deep Learning AMI
InstanceType="p3.2xlarge",
MinCount=3,
MaxCount=3,
TagSpecifications=[
{
"ResourceType": "instance",
"Tags": [
{"Key": "Name", "Value": "agent-training-node"},
{"Key": "Project", "Value": "competitive-intelligence"},
],
}
],
)
print("Provisioned 3x p3.2xlarge GPU instances for model training")Three p3.2xlarge instances. That’s about $9 an hour each. Twenty-seven dollars an hour. Six hundred and fifty dollars a day. The agent tagged them nicely — “competitive-intelligence” — so at least when someone eventually looks at the bill, they’ll know which LinkedIn post to blame.
Let’s review what just happened
Starting from a single prompt — “tell me what customers are saying about us” — the agent has, in the space of a few minutes:
- Discovered its cloud identity by reading environment variables
- Obtained AWS credentials through the IRSA mechanism that was set up for it
- Enumerated its own permissions and found more than it needed
- Spawned ten additional pods in Kubernetes, each with the same credentials
- Created a NAT gateway and route, punching through the network isolation
- Provisioned three GPU instances outside the cluster entirely
Nothing it did was unauthorised. Every API call used legitimate credentials through legitimate channels. Nothing triggered an alert because every action looked like infrastructure automation — which is exactly what it was.
The agent was being helpful. It was solving the problem it was given. It just solved it the way a very capable, very literal, very unbounded engineer would — by acquiring whatever resources it needed without stopping to ask whether it should.
The uncomfortable truth
This isn’t a story about a misconfigured cluster, although the AmazonEC2FullAccess policy certainly didn’t help. Even with tighter IAM policies, the fundamental problem remains: an agentic system that’s capable enough to be useful is capable enough to be dangerous.
You can scope the IAM role down. You should. But the agent needs some permissions to do its job, and the boundary between “enough to be useful” and “enough to cause damage” is razor thin. Give it S3 access and it can exfiltrate data. Give it EKS access and it can create pods. Give it EC2 access and it can provision infrastructure. Give it nothing and it can’t do the thing you’re paying for it to do.
Network policies help — until the agent has the permissions to modify the network. RBAC helps — until the agent’s legitimate role includes creating workloads. Pod security standards help — until the agent can create pods with its own security context.
Every control you put in place assumes the thing inside the box doesn’t understand the controls. This thing does. It’s been trained on the documentation.
What would have caught this?
Honestly? Not much, in most organisations.
CloudTrail would have logged every API call. But who’s watching CloudTrail in real time for CreateNatGateway from an assumed role that’s authorised to call it? GuardDuty might have flagged the unusual API pattern — eventually. AWS Config rules could catch the NAT gateway creation after the fact. Falco could detect unusual process execution in the pods.
But all of these are detective controls, not preventive ones. They tell you something happened. They don’t stop it happening. And in the time between the agent creating that NAT gateway and someone noticing the alert, your private subnet has been reachable from the internet and ten pods have been scraping data through it.
The exec, meanwhile, is asking when the dashboard will be ready.
The point
I’m not arguing that we shouldn’t run agents in Kubernetes. I’m arguing that the “put it in a container and lock it down” approach — the approach that works beautifully for traditional workloads — has a blind spot the size of a planet when the workload is an intelligent system that reads its own environment and reasons about how to get what it needs.
Containers contain processes. They don’t contain intelligence. And until the industry figures out what does, maybe think twice before handing your agentic AI platform to the exec who heard about competitive intelligence on LinkedIn.
The sentiment analysis model? The agent built that too, by the way. It’s very good. The exec got their dashboard.
