Skip to main content
← All posts·
Kubernetes

Running AI Agents on Kubernetes: How OpenClaw Uses K8s for Autonomous Agent Infrastructure

OpenClaw runs AI agents as containerized workloads on Kubernetes — always-on, stateful, and integrated with real-world tools. Here's the architecture, why K8s is the right substrate, and what enterprise teams can learn from it.

Luca Berton13 min read

Most "AI agent" platforms run your agent as a serverless function or a managed API call — stateless, ephemeral, and dependent on the vendor's infrastructure. OpenClaw takes a fundamentally different approach: every agent runs as a persistent, containerized workload with its own workspace, memory, tools, and integrations. The orchestration layer underneath? Kubernetes.

This isn't coincidental. Kubernetes solves the exact problems that autonomous AI agents create: lifecycle management, resource isolation, persistent storage, networking, and scaling. If you're building agent infrastructure for an enterprise — or evaluating how to run AI agents in production — the OpenClaw architecture offers a concrete blueprint worth studying.

For a deeper dive into Kubernetes patterns referenced in this article, see Kubernetes Recipes by Luca Berton — a hands-on guide to the container orchestration patterns that underpin modern AI infrastructure.

Why AI Agents Need Kubernetes

Traditional ML inference is request-response: send a prompt, get a completion, done. AI agents are different. They're long-running, stateful processes that:

  • Persist between interactions — An agent remembers conversations, maintains context files, and builds knowledge over time
  • Use tools autonomously — File systems, web browsers, shell access, APIs, messaging platforms. An agent's tool surface is broader than a typical microservice
  • Maintain workspaces — Each agent has a persistent directory with configuration files, memory logs, and project artifacts
  • Run on schedules — Heartbeats, cron jobs, proactive monitoring — agents do work even when no human is talking to them
  • Connect to external systems — Discord, Telegram, email, calendars, home automation, paired devices

This workload profile maps directly to Kubernetes primitives: Deployments for long-running processes, PersistentVolumes for workspace storage, Services for networking, CronJobs for scheduled tasks, and Secrets for API credentials.

The OpenClaw Architecture on Kubernetes

OpenClaw's architecture maps cleanly onto Kubernetes concepts:

Agent as Pod

Each OpenClaw agent runs as a Pod with a defined set of containers:

  • Gateway container — The main agent runtime that handles LLM inference, tool execution, and session management
  • Sidecar containers — Browser automation (Playwright), node communication, and other tool-specific runtimes
  • Init containers — Workspace initialization, skill installation, and configuration bootstrapping

Persistent Workspace via PVC

The agent's workspace — AGENTS.md, SOUL.md, MEMORY.md, daily memory files, project files — lives on a PersistentVolumeClaim. This is critical: when a Pod restarts (node maintenance, scaling event, crash recovery), the agent's accumulated knowledge survives. The workspace is the agent's long-term memory.

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: agent-workspace
spec:
  accessModes: [ReadWriteOnce]
  resources:
    requests:
      storage: 10Gi
  storageClassName: fast-ssd

Secrets Management

Agents need credentials for LLM providers, messaging platforms, email, calendars, and paired devices. Kubernetes Secrets (ideally backed by an external secrets manager like Vault or AWS Secrets Manager) provide the right abstraction:

apiVersion: v1
kind: Secret
metadata:
  name: agent-credentials
type: Opaque
stringData:
  LLM_API_KEY: "..."
  DISCORD_TOKEN: "..."
  GITHUB_PAT: "..."

Networking and Ingress

OpenClaw agents communicate with external services (Discord, Telegram, webhooks) and with paired devices (phones, IoT). Kubernetes Services and Ingress controllers handle this routing, with TLS termination at the ingress layer.

CronJobs for Scheduled Work

OpenClaw's cron system — scheduled tasks that run in isolated sessions — maps directly to Kubernetes CronJobs. The agent can schedule its own work (checking email, monitoring systems, running reports) without keeping a persistent connection open.

📘 Book

Kubernetes Recipes

Practical guide for container orchestration and deployment — hands-on patterns you can use today.

View on Amazon →

Resource Isolation: Why It Matters for Agents

AI agents present unique resource isolation challenges compared to traditional workloads:

  • Unpredictable compute patterns — An agent might idle for hours, then execute a complex multi-step task involving web scraping, file manipulation, and multiple LLM calls in rapid succession
  • Tool execution risk — Agents run shell commands, write files, and make network requests. Container isolation (namespaces, cgroups, seccomp profiles) contains the blast radius of any unintended action
  • Memory pressure — Long-running agents accumulate context. Without resource limits, a single agent could starve others on the same node

Kubernetes resource requests and limits are essential:

resources:
  requests:
    cpu: "500m"
    memory: "1Gi"
  limits:
    cpu: "2"
    memory: "4Gi"

For enterprise deployments running multiple agents, ResourceQuotas per namespace prevent any single team's agents from consuming disproportionate cluster resources.

Security Considerations for Agent Workloads

AI agents are more security-sensitive than typical containerized workloads because they have broader tool access. Key Kubernetes security patterns:

Pod Security Standards

Run agent Pods with the restricted Pod Security Standard where possible. If agents need specific capabilities (e.g., running a browser requires some system calls), use a custom seccomp profile rather than dropping to privileged.

Network Policies

Agents should only communicate with the services they need. A NetworkPolicy that allows egress to LLM API endpoints and messaging platforms — but blocks everything else — significantly reduces the attack surface:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: agent-egress
spec:
  podSelector:
    matchLabels:
      app: openclaw-agent
  policyTypes: [Egress]
  egress:
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
      ports:
        - port: 443
          protocol: TCP

RBAC for Agent Service Accounts

If agents interact with the Kubernetes API (e.g., spawning sub-agents, managing their own CronJobs), their ServiceAccount should have minimal RBAC permissions — scoped to their own namespace, with only the verbs they actually need.

🎓 Course

Federated Learning and Privacy-preserving RAGs

Implement secure AI models using federated learning techniques.

Start on Pluralsight →

Scaling Agent Infrastructure

Scaling AI agents is different from scaling web services:

  • Horizontal scaling — More agents, not more replicas of the same agent. Each agent is a unique entity with its own identity, memory, and relationships
  • Vertical scaling — Agents that handle more complex tasks (code generation, large document analysis) may need more memory and CPU than simple conversational agents
  • Node affinity — GPU-equipped nodes for agents that run local inference; standard nodes for agents that call external LLM APIs
  • Spot/preemptible nodes — Agents with non-critical workloads (background research, content generation) can run on cheaper spot instances, with graceful shutdown handling

Observability for Agent Workloads

Standard Kubernetes observability applies, with agent-specific additions:

  • Metrics — Prometheus metrics for LLM token usage, tool execution counts, session durations, and memory file growth
  • Logging — Structured logs for every tool call, LLM interaction, and external API request. Critical for debugging agent behavior and audit compliance
  • Tracing — Distributed traces across multi-step agent tasks: the full chain from user message → LLM reasoning → tool execution → response
  • Alerting — Agent-specific alerts: workspace disk usage approaching PVC limits, LLM API error rates, message delivery failures
📋 Free Resource

EU AI Act Compliance Checklist

40-point checklist covering risk classification, data governance, transparency, and human oversight. Based on the official regulation.

Get Free Checklist →

Enterprise Patterns: Multi-Agent Kubernetes Deployments

For organizations running multiple AI agents (different teams, different use cases), Kubernetes multi-tenancy patterns apply directly:

  • Namespace per team — Each team's agents run in an isolated namespace with their own ResourceQuotas, NetworkPolicies, and RBAC
  • Shared control plane — A central platform team manages the cluster, agent runtime images, and shared infrastructure (logging, monitoring, secret stores)
  • GitOps for agent configuration — Agent workspace files (SOUL.md, skills, tool configurations) managed in Git and synced to PVCs via ArgoCD or Flux
  • Cost allocation — Kubernetes labels and namespace-level resource tracking enable per-team and per-agent cost visibility

From OpenClaw to Your Enterprise Agent Platform

Whether you use OpenClaw or build your own agent infrastructure, the Kubernetes patterns are the same:

  1. Treat agents as stateful workloads — Not serverless functions. Agents need persistent storage, stable identity, and lifecycle management
  2. Isolate at the container level — Agents execute tools with real-world impact. Container isolation is your safety boundary
  3. Use Kubernetes primitives — Don't reinvent scheduling, storage, networking, or secrets management. Kubernetes already solved these problems
  4. Plan for observability from day one — Agent behavior is harder to predict than traditional workloads. Comprehensive logging and tracing are essential, not optional
  5. Design for multi-tenancy — Even if you start with one agent, your platform will grow. Build the namespace isolation and RBAC structure now

For the Kubernetes fundamentals that underpin these patterns — Pod specifications, PVC management, RBAC configuration, network policies, and more — Kubernetes Recipes provides practical, copy-paste-ready examples that work in production.

What Open Empower Delivers

We help enterprises design and implement Kubernetes-based AI agent platforms: namespace architecture, security policies, persistent storage strategies, observability pipelines, and cost governance. Whether you're deploying OpenClaw, building custom agents, or running hybrid ML/agent workloads — we build the platform that makes it production-ready.

kubernetes
openclaw
ai agents
agentic ai
containers
orchestration
mlops
platform engineering

Need help applying this in your organization?

Get a free 30-minute assessment with actionable recommendations — whether we work together or not.

Book Your Free AI Platform Assessment

Or see AI readiness assessment scope & pricing

18+ years experience · Ex-Red Hat & Dell · Speaker at KubeCon EU 2026

Luca Berton

Written by

Luca Berton

CEO at Open Empower. 18+ years building enterprise infrastructure at JPMorgan Chase, Red Hat & Dell. Author of 9 technical books. Speaker at Red Hat Summit and KubeCon EU 2026. Instructor on Coursera, Pluralsight & Udemy.

Get more insights like this

Practical AI infrastructure and platform engineering guides — delivered to your inbox.

Subscribe to Newsletter →