justinsb/kube-agents

β˜… 0Forks 0GitHub β†—Compare

README

🧭 kube-agents β€” The Kubernetes Agentic Harness

Stop driving your clusters. Start delegating them.

kube-agents replaces the traditional imperative DevOps presentation layer β€” kubectl, gcloud, the Google Cloud Console β€” with autonomous, proactive AI agents that manage your Kubernetes/GKE infrastructure, enforce multi-tenant governance, and continuously audit security posture. Instead of you reacting to pages and typing commands, a Platform Agent watches your fleet around the clock, opens pull requests with fixes, and reports to you in chat.

Traditional Ops With kube-agents
Reactive, manual toil (kubectl + runbooks) Proactive, intent-driven operations
Drift discovered during incidents Scheduled compliance & blueprint audits (autonomous watchdogs)
Hand-rolled RBAC and tenancy reviews Automated RBAC & boundary enforcement, credential isolation by design
Patch Tuesdays and CVE spreadsheets Daily vulnerability & patch scans with staggered rollout orchestration
One human, one terminal ChatOps with the agent over Google Chat & Slack

πŸ“— Full documentation: gke-labs.github.io/kube-agents


⚑ Try it now

The fastest way in: clone this repository into your agent harness's workspace and delegate the setup itself to an agent:

"Using kube-agents/INSTALL.md provision k8s agentic harness and create platform agent"

INSTALL.md is written so that an AI agent with file access and shell tools can follow it end-to-end β€” and the same guide works step-by-step for humans.

Prefer the scripted path? From an authenticated gcloud:

cd k8s-operator
make gcp-provision

This runs the staged, idempotent provisioning scripts end to end β€” from GKE cluster creation through IAM and chat integration to the optional inference stack. make gcp-teardown reverses it. See the quick start for the walkthrough, or INSTALL.md for manual and local-development paths.


πŸ“– What it is

The harness runs co-located agents in a single operator-deployed pod: the Chat Agent β€” the conversational front door that receives every chat message and delegates work over a shared kanban board β€” the Platform Agent β€” the master custodian and agent architect that manages the GKE infrastructure lifecycle, establishes multi-tenancy boundaries, and enforces fleet-wide compliance β€” and a Cluster Agent per managed cluster, a single-cluster SRE persona the Platform Agent scaffolds from the agents/cluster/ template for runtime operations and workload debugging, with read-only access to the cluster it watches. The Platform Agent is driven by:

  • 🧬 A persona β€” agents/platform/SOUL.md defines its identity, its Automation First rule (no manual cluster mutations; changes flow through declarative, PR-based workflows), and its Least Privilege constraint.
  • πŸ“š Governance playbooks β€” SOPs in agents/platform/governance/ covering blueprint sync, compliance audits, cost analysis, capacity orchestration, security patch orchestration, and lifecycle management.
  • πŸ› οΈ Skills β€” task-focused SKILL.md bundles under agents/platform/skills/: cluster creation, app onboarding, cost analysis, backup & DR, and manifest generation. Single-cluster runtime skills β€” workload troubleshooting, observability, autoscaling, storage β€” belong to the Cluster Agent in agents/cluster/skills/. See the skill catalog.
  • ⏰ Autonomous watchdogs β€” cron-driven governance jobs in agents/platform/cron/jobs.json that keep the fleet honest without human prompting. See proactive autonomy.

The runtime is built on the Hermes agent framework and wires in MCP servers for platform control and GKE's hosted MCP endpoint, so the agent speaks to your clusters through structured tools rather than raw shell access.


πŸ›‘οΈ Governance & isolation

kube-agents is designed for enterprise fleets where agents must be powerful and provably contained:

  • Least-privilege RBAC β€” the agent's Kubernetes identity is read-only and cannot read Secrets.
  • Credential isolation β€” the agent sandbox container never receives API keys or tokens; an Envoy credential-proxy sidecar injects them at the network boundary.
  • Kernel-level sandboxing β€” agent workloads can run under a gVisor RuntimeClass (GKE Sandbox).
  • GitOps-only mutations β€” infrastructure changes are proposed as pull requests for human review.

Exactly what is enforced on which plane β€” Kubernetes RBAC, GCP IAM, and the GitOps path each answer differently β€” is set out in Security & IAM. Read that before granting the agent access to a production project.


πŸ—οΈ Architecture

flowchart TB
    subgraph agent["🧠 Control Plane β€” Agent Layer"]
        SOUL["SOUL.md persona<br/>+ governance SOPs"]
        SKILLS["Skills<br/>(agents/platform/skills)"]
        CRON["Scheduled watchdogs<br/>(cron/jobs.json)"]
        PA["Platform Agent workspace<br/>(agents/platform)"]
        SOUL --> PA
        SKILLS --> PA
        CRON --> PA
    end

    subgraph cluster["☸️ Cluster Plane β€” Kubernetes Layer"]
        OP["k8s-operator<br/>(Go / Kubebuilder)"]
        CRD["PlatformAgent CRD<br/>kubeagents.x-k8s.io/v1alpha1"]
        POD["Agent pod: gVisor sandbox<br/>+ Envoy credential proxy<br/>+ Fluent Bit + event watcher"]
        RBAC["RBAC isolation boundaries<br/>(read-only view + explorer)"]
        OP -->|reconciles| CRD
        CRD --> POD
        OP --> RBAC
    end

    subgraph integration["πŸ”€ Integration & Routing Layer"]
        LLM["LiteLLM Gateway<br/>Gemini Β· OpenAI Β· Anthropic"]
        CHAT["Messaging bridges<br/>Google Chat (Pub/Sub) Β· Slack (Socket Mode)"]
        GH["Minty β€” GitHub App<br/>token minter (KMS)"]
    end

    PA -.runs inside.-> POD
    POD --> LLM
    CHAT <--> POD
    POD -->|PR-based changes| GH
Loading

Walkthrough: Architecture. The k8s-operator/ reconciles PlatformAgent custom resources into the sandboxed agent pod, its sidecars, per-agent ServiceAccounts with Workload Identity, read-only RBAC, and Services.

Looking for the end-state design? docs/architecture/ specifies a three-tier, fully read-only agent model that this repository is converging toward. It describes the target, not what ships today.


🀝 Contributing

Contributions are welcome. See docs/contributing.md for CLA requirements and the contributing guide for PR hygiene, commit conventions, and the local checks CI enforces. Repository conventions for AI coding agents are in AGENTS.md.

Disclaimer

This is not an officially supported Google product.

This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.

Contributors

bradhoekstramplakhtiydshnayderDzianis-Harbatsenkadependabot[bot]mtaufentoshiowangkirill-kosovjayantidstalhaalieLeontevshalinibhatiaplsergiu1spiridonfatoshotihaoxuwadrianchungmastersingh24AntonTyblapis2002justinsbbnayloradamparcofkc1e100mateuszklinowskierain

Issues