Compute Cluster Forge is a modular infrastructure-as-code (IaC) framework that leverages Terraform to provision secure, auto-healing, and cost-optimized AI/ML compute clusters. While currently configured for Google Cloud Platform (GCP), the architecture is designed for seamless extension to AWS and Azure.
This implementation adheres to industry-standard DevOps practices, moving beyond static scripting toward a dynamic, policy-driven infrastructure:
- π Zero-Trust Security (Google IAP): Implements a strict security model where all public ingress is disabled. Management access (SSH/Web) is limited to authorized users via Google Identity-Aware Proxy (IAP) TCP forwarding.
- π₯· Proactive Instance Remediation: Implements specialized recovery logic in the health probe to detect
ZONE_RESOURCE_POOL_EXHAUSTEDerrors. The system automatically purges failing instances and forces the MIG to 'hop' to alternative zones within the region to find available hardware capacity. - π©Ή Auto-Healing Spot VM Orchestration: Leverages Google's cost-effective Spot instances while utilizing Managed Instance Groups (MIGs) with the
ANYdistribution shape andallow_changing_zonepolicy to provide automated cross-zone self-healing. - π§ VRAM-Optimized AI Features: Provides turn-key integration for LLM runners (Ollama) and pre-configured model tiers (Gemma4 26B, 31B, E2B). Models are surgically matched to hardware profiles (e.g., L4 GPUs) to ensure 100% VRAM offloading and near-instantaneous inference.
- πΊοΈ Dynamic Infrastructure Discovery: Implements automated resource discovery. The system queries regional metadata to identify zones supporting specific hardware requirements and restricts deployment to compatible areas.
- π» Cross-OS SSH Orchestration: Includes an automated PowerShell bridge that natively configures Windows VS Code Remote-SSH profiles with secure IAP ProxyCommands, ensuring seamless connectivity regardless of zone-hopping events.
/compute-cluster-forge
βββ /modules # Infrastructure-as-Code (IaC) Modules
β βββ /aws # [Planned] AWS Provider Module
β βββ /azure # [Planned] Azure Provider Module
β βββ /gcp # Active GCP Provider Module
β βββ main.tf # Resource Orchestration (MIG, Firewalls, APIs)
β βββ variables.tf # Input Schema & Secure Defaults
β βββ outputs.tf # Validated Resource References
β βββ startup-script.sh # Native Bootstrapping Logic
βββ run.sh # CLI Orchestration & Execution Wrapper
βββ /templates # Environment & Hardware Profiles
βββ training.tfvars # High-compute Base Profile
βββ inference.tfvars # Low-latency Base Profile
βββ spike.tfvars # General-purpose Burst Profile
βββ gpu-l4.tfvars # Hardware Layer: NVIDIA L4 GPU / G2 Family
Ensure the following toolchains are available in your environment:
- Terraform (v1.0+)
- Google Cloud CLI (
gcloud) - PowerShell 7+ (
pwsh) (Windows/WSL users only. Required for the orchestrator to automatically bridge WSL and natively configure your Windows VS Code Remote-SSH profiles).
Authenticate your local machine to generate Application Default Credentials (ADC):
gcloud auth application-default loginUtilize the run.sh orchestrator to provision clusters. The interface supports combining environment templates with hardware layers and runtime variable overrides.
Example: Provisioning an AI-Ready Cluster with NVIDIA L4 & Gemma4-26B:
./run.sh --cloud gcp --type spike --layer gpu-l4 --feature ollama --feature gemma4-26b --var project_id=[PROJECT_ID] --action applyWhat happens behind the scenes?
- Isolation: Terraform initializes within an isolated workspace.
- Discovery: Executes dynamic zone discovery based on hardware requirements.
- Provisioning: Deploys IAM policies, networking, storage, and compute resources.
- Feature Injection: Dynamically generates startup scripts to install requested software (e.g., Ollama runner) and pull VRAM-optimized models.
- Verification: Automated status verification monitors the Instance Group until version targets are reached, health checks pass, and features are initialized.
- Reporting: Generates instance-specific connection strings for secure access.
You must provide the project id at runtime using one of the following methods:
Set the standard Terraform environment variable to automatically inject the project id into all forge commands within your current shell session:
export TF_VAR_project_id="your-gcp-project-id"
./run.sh --cloud gcp --type spike --action applyPass the project id directly using the --var flag:
./run.sh --cloud gcp --type spike --var project_id="your-gcp-project-id" --action applyInternal services must be accessed through the established IAP tunnel.
To ensure accessibility via IAP TCP forwarding, all applications (e.g., Jupyter Lab, Tensorboard) MUST bind to 0.0.0.0. Applications listening exclusively on 127.0.0.1 will be unreachable.
- β
Jupyter Lab:
jupyter lab --ip=0.0.0.0 --port=8888 --no-browser - β
HTTP Server:
python3 -m http.server 8080 --bind 0.0.0.0
Once the cluster is live, run.sh provides dynamically resolved connection commands:
Secure SSH Access:
gcloud compute ssh [INSTANCE_NAME] --tunnel-through-iap --project=[PROJECT_ID] --zone=[ZONE]IAP Port Forwarding (Remote 8888 β Local 8888):
gcloud compute start-iap-tunnel [INSTANCE_NAME] 8888 --local-host-port=localhost:8888 --project=[PROJECT_ID] --zone=[ZONE]Because the architecture utilizes Proactive Instance Remediation, an instance may be recreated in a different zone if the original zone experiences resource exhaustion. Since the VS Code ProxyCommand is zone-dependent, a "Zone-Hop" will cause existing SSH config entries to become obsolete. To restore connectivity, simply re-run the apply command (or the automated PowerShell bridge) to surgically update your local .ssh/config with the new instance location.
To prevent unnecessary billing, decommission all provisioned resources upon task completion:
./run.sh --cloud gcp --type spike --layer gpu-l4 --var project_id=[PROJECT_ID] --action destroyπ Safety Note: This operation performs a forceful deletion of all regional resources, including the stateful GCS bucket. Ensure all datasets and trained models are synchronized to local storage prior to execution.