Stop driving your clusters. Start delegating them.
kube-agents replaces the traditional imperative DevOps presentation layer β kubectl, gcloud, the Google Cloud Console β with autonomous, proactive AI agents that manage your Kubernetes/GKE infrastructure, enforce multi-tenant governance, and continuously audit security posture. Instead of you reacting to pages and typing commands, a Platform Agent watches your fleet around the clock, opens pull requests with fixes, and reports to you in chat.
| Traditional Ops | With kube-agents |
|---|---|
Reactive, manual toil (kubectl + runbooks) |
Proactive, intent-driven operations |
| Drift discovered during incidents | Scheduled compliance & blueprint audits (autonomous watchdogs) |
| Hand-rolled RBAC and tenancy reviews | Automated RBAC & boundary enforcement, credential isolation by design |
| Patch Tuesdays and CVE spreadsheets | Daily vulnerability & patch scans with staggered rollout orchestration |
| One human, one terminal | ChatOps with the agent over Google Chat & Slack |
π Full documentation: gke-labs.github.io/kube-agents
The fastest way in: clone this repository into your agent harness's workspace and delegate the setup itself to an agent:
"Using kube-agents/INSTALL.md provision k8s agentic harness and create platform agent"
INSTALL.md is written so that an AI agent with file access and shell tools can follow it end-to-end β and the same guide works step-by-step for humans.
Prefer the scripted path? From an authenticated gcloud:
cd k8s-operator
make gcp-provisionThis runs the staged, idempotent provisioning scripts end to end β from GKE cluster creation through IAM and chat integration to the optional inference stack. make gcp-teardown reverses it. See the quick start for the walkthrough, or INSTALL.md for manual and local-development paths.
The harness runs co-located agents in a single operator-deployed pod: the Chat Agent β the conversational front door that receives every chat message and delegates work over a shared kanban board β the Platform Agent β the master custodian and agent architect that manages the GKE infrastructure lifecycle, establishes multi-tenancy boundaries, and enforces fleet-wide compliance β and a Cluster Agent per managed cluster, a single-cluster SRE persona the Platform Agent scaffolds from the agents/cluster/ template for runtime operations and workload debugging, with read-only access to the cluster it watches. The Platform Agent is driven by:
- 𧬠A persona β
agents/platform/SOUL.mddefines its identity, its Automation First rule (no manual cluster mutations; changes flow through declarative, PR-based workflows), and its Least Privilege constraint. - π Governance playbooks β SOPs in
agents/platform/governance/covering blueprint sync, compliance audits, cost analysis, capacity orchestration, security patch orchestration, and lifecycle management. - π οΈ Skills β task-focused
SKILL.mdbundles underagents/platform/skills/: cluster creation, app onboarding, cost analysis, backup & DR, and manifest generation. Single-cluster runtime skills β workload troubleshooting, observability, autoscaling, storage β belong to the Cluster Agent inagents/cluster/skills/. See the skill catalog. - β° Autonomous watchdogs β cron-driven governance jobs in
agents/platform/cron/jobs.jsonthat keep the fleet honest without human prompting. See proactive autonomy.
The runtime is built on the Hermes agent framework and wires in MCP servers for platform control and GKE's hosted MCP endpoint, so the agent speaks to your clusters through structured tools rather than raw shell access.
kube-agents is designed for enterprise fleets where agents must be powerful and provably contained:
- Least-privilege RBAC β the agent's Kubernetes identity is read-only and cannot read Secrets.
- Credential isolation β the agent sandbox container never receives API keys or tokens; an Envoy credential-proxy sidecar injects them at the network boundary.
- Kernel-level sandboxing β agent workloads can run under a gVisor RuntimeClass (GKE Sandbox).
- GitOps-only mutations β infrastructure changes are proposed as pull requests for human review.
Exactly what is enforced on which plane β Kubernetes RBAC, GCP IAM, and the GitOps path each answer differently β is set out in Security & IAM. Read that before granting the agent access to a production project.
flowchart TB
subgraph agent["π§ Control Plane β Agent Layer"]
SOUL["SOUL.md persona<br/>+ governance SOPs"]
SKILLS["Skills<br/>(agents/platform/skills)"]
CRON["Scheduled watchdogs<br/>(cron/jobs.json)"]
PA["Platform Agent workspace<br/>(agents/platform)"]
SOUL --> PA
SKILLS --> PA
CRON --> PA
end
subgraph cluster["βΈοΈ Cluster Plane β Kubernetes Layer"]
OP["k8s-operator<br/>(Go / Kubebuilder)"]
CRD["PlatformAgent CRD<br/>kubeagents.x-k8s.io/v1alpha1"]
POD["Agent pod: gVisor sandbox<br/>+ Envoy credential proxy<br/>+ Fluent Bit + event watcher"]
RBAC["RBAC isolation boundaries<br/>(read-only view + explorer)"]
OP -->|reconciles| CRD
CRD --> POD
OP --> RBAC
end
subgraph integration["π Integration & Routing Layer"]
LLM["LiteLLM Gateway<br/>Gemini Β· OpenAI Β· Anthropic"]
CHAT["Messaging bridges<br/>Google Chat (Pub/Sub) Β· Slack (Socket Mode)"]
GH["Minty β GitHub App<br/>token minter (KMS)"]
end
PA -.runs inside.-> POD
POD --> LLM
CHAT <--> POD
POD -->|PR-based changes| GH
Walkthrough: Architecture. The k8s-operator/ reconciles PlatformAgent custom resources into the sandboxed agent pod, its sidecars, per-agent ServiceAccounts with Workload Identity, read-only RBAC, and Services.
Looking for the end-state design?
docs/architecture/specifies a three-tier, fully read-only agent model that this repository is converging toward. It describes the target, not what ships today.
Contributions are welcome. See docs/contributing.md for CLA requirements and the contributing guide for PR hygiene, commit conventions, and the local checks CI enforces. Repository conventions for AI coding agents are in AGENTS.md.
This is not an officially supported Google product.
This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.