latentwill/runpod-skill

Claude Code skill for RunPod GPU infrastructure management

★ 0Forks 0GitHub ↗Compare

README

runpod-skill

A Claude Code skill for deploying and managing GPU/CPU infrastructure on RunPod.

Built for any AI agent that needs to programmatically launch pods, manage storage, and control compute resources — not tied to a specific workflow or use case.

What it covers

  • REST API v2 — built against https://api.runpod.io/v2 (v1 retires December 1, 2026)
  • Pod lifecycle — create, list, start, stop, restart, update, terminate via the consolidated /action endpoint
  • GPU selection — catalog endpoints for availability + full GPU table with IDs, VRAM, and fallback chains by use case
  • Spot instances — podRentInterruptable via GraphQL with bid pricing
  • Network volumes — persistent storage independent of pods (CRUD + attach via mounts)
  • Templates — reusable pod configurations via REST API
  • Connectivity — SSH (ssh.proxy/ssh.direct blocks), HTTP proxy, TCP ports with timeout warnings
  • Billing — cost queries for pods, serverless, network volumes, and endpoints
  • CLI — runpodctl reference for common operations
  • Error handling — RFC 9457 problem format, retry patterns for GPU stock-outs, zero-GPU restarts, volume mismatches

What it gets right that others get wrong

Common mistake This skill
Stale v1 URL (rest.runpod.io/v1) Correct: api.runpod.io/v2 (REST v2 lives here; v1 is deprecated)
Flat v1 create body (gpuTypeIds, imageName, volumeInGb) Correct: nested v2 body (gpu: {id, count}, image, mounts)
Separate /stop /start /restart endpoints Correct: single POST /v2/pods/{id}/action with {"action": ...}
Unwrapped v1 list responses Correct: v2 wraps lists ({"pods": [...]}, {"networkVolumes": [...]})
GraphQL for availability Correct: REST catalog (/v2/catalog/gpus)
region: "us-west-1" (doesn't exist) Correct: dataCenterIds: ["US-TX-3"]
env format same for both APIs Documents the difference: REST = object, GraphQL = array of {key, value}
gpuTypeId vs gpuTypeIds confusion Flags singular (GraphQL) vs nested gpu: {id, count} (REST v2)
Container disk is persistent (it isn't) Explicit ephemeral/persistent/network storage model

Install

Copy the SKILL.md file into your Claude Code skills directory:

# If you have a skills directory configured
cp SKILL.md ~/.claude/skills/runpod/SKILL.md

# Or place it alongside your project
cp SKILL.md .claude/skills/runpod/SKILL.md

The skill activates when prompts mention RunPod, GPU cloud, launching pods, or related terms.

Prerequisites

  • A RunPod account with API access
  • API key set as RUNPOD_API_KEY environment variable

API coverage

The skill standardizes on the REST API v2 (api.runpod.io/v2) for all CRUD operations and documents GraphQL (api.runpod.io/graphql) for two things REST v2 can't do:

  1. Runtime GPU utilization metrics (GPU utilization, container CPU/memory)
  2. Spot instance deployment (podRentInterruptable)

GPU availability queries — previously GraphQL-only — now use the REST v2 catalog (GET /v2/catalog/gpus).

License

MIT

Contributors

latentwill

Issues