Complete infrastructure-as-code setup for deploying a production-ready Transaction API on Google Kubernetes Engine (GKE) with comprehensive monitoring and observability.
This repository contains Terraform modules and Helm charts for deploying:
- GKE Cluster - Google Kubernetes Engine cluster with autoscaling
- Transaction API - RESTful API for transaction processing
- PostgreSQL Database - Persistent data storage
- Monitoring Stack - Prometheus, Grafana, AlertManager
- Observability - Metrics, alerts, dashboards, and SLO tracking
.
βββ app/
β βββ transaction-api/ # Transaction API application deployment
βββ infrastructure/
β βββ cluster/ # GKE cluster infrastructure
β βββ image-repo/ # Docker image artifact registry
β βββ monitoring/ # Prometheus monitoring stack
βββ modules/
β βββ cluster/ # Reusable GKE cluster module
β βββ image-repo/ # Reusable image repository module
β βββ postgresql/ # PostgreSQL Helm chart module
β βββ prometheus/ # Prometheus stack module
β βββ transaction-api/ # Transaction API Helm chart module
βββ docs/
β βββ Deployment/ # Deployment guides
β βββ Monitoring/ # Monitoring setup and guides
β βββ RunBooks/ # Operational runbooks
βββ README.md # This file
| Document | Description |
|---|---|
| DEPLOYMENT.md | Complete step-by-step deployment guide with prerequisites, deployment order, verification steps, and troubleshooting |
Key Topics:
- Prerequisites and tool setup
- Infrastructure deployment (image-repo, cluster, monitoring)
- Application deployment (Transaction API + PostgreSQL)
- Verification procedures
- Rollback strategies
- Cleanup/teardown instructions
| Document | Description |
|---|---|
| README.md | Monitoring overview, architecture, and quick start guide |
| transaction-api-monitoring-guide.md | Complete monitoring implementation guide with SLOs, alerts, and dashboards |
| QUICK_REFERENCE.md | Quick reference card for SLOs, metrics queries, alerts, and troubleshooting |
Key Topics:
- Prometheus metrics collection
- Grafana dashboard setup
- Service Level Objectives (SLOs)
- Alert rules and thresholds
- Error budget tracking
- Application instrumentation examples (Python/Node.js)
| Runbook | Alert | Description |
|---|---|---|
| README.md | - | Runbooks overview and quick reference |
| high-error-rate.md | TransactionAPIHighErrorRate | Error rate > 1% - deployment issues, database problems, resource exhaustion |
| service-down.md | TransactionAPIDown | Complete service outage - pod crashes, scaling issues, network problems |
| database-errors.md | DatabaseConnectionFailures | DB error rate > 0.1% - connection pool, slow queries, network issues |
| high-latency.md | TransactionAPIHighLatency | P95 > 200ms - CPU pressure, slow queries, external API delays |
Each runbook includes:
- Alert description and thresholds
- Quick diagnosis steps with PromQL queries
- Common causes and solutions
- Step-by-step investigation procedures
- Verification steps
- Escalation procedures
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β GCP Project β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββ β
β β Artifact β β GCS Bucket β β GKE Cluster β β
β β Registry β β (TF State) β β β β
β ββββββββββββββββ ββββββββββββββββ ββββββββ¬ββββββββ β
β β β
β βββββββββββββββββββββββββββββββββββββββββ β
β β β
β ββββββΌβββββββββββββββββββββββββββββββββββββββββββββββ β
β β GKE Cluster Namespaces β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββ€ β
β β ββββββββββββββββ ββββββββββββββββ β β
β β βtransactions β β monitoring β β β
β β β β β β β β
β β ββ’ Trans API β ββ’ Prometheus β β β
β β ββ’ PostgreSQL β ββ’ Grafana β β β
β β β β ββ’ AlertManager β β β
β β ββββββββββββββββ ββββββββββββββββ β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Transaction API (App)
β (exposes /metrics)
Prometheus (scrapes metrics)
β
ββββββ΄βββββ
β β
Grafana AlertManager
(visualize) (notify)
β
Multi-zone GKE cluster with autoscaling
β
Managed node pools with auto-repair and auto-upgrade
β
Artifact Registry for Docker images
β
Infrastructure as Code - Complete Terraform setup
β
Makefile automation - Simple deployment commands
β
Horizontal Pod Autoscaling (HPA)
β
Pod Disruption Budgets (PDB)
β
Health checks and liveness probes
β
Resource limits and requests
β
Pod anti-affinity for high availability
| Category | Technology |
|---|---|
| Cloud Provider | Google Cloud Platform (GCP) |
| Container Orchestration | Google Kubernetes Engine (GKE) |
| Infrastructure as Code | Terraform |
| Package Management | Helm |
| Monitoring | Prometheus, Grafana, AlertManager |
| Database | PostgreSQL |
| Programming | Go (Transaction API) |
- Create feature branch
- Make changes
- Test locally
- Run
terraform planto preview changes - Submit pull request
- Deploy to staging first
- Verify in staging
- Deploy to production
- Use Terraform formatting:
terraform fmt -recursive - Validate Terraform:
terraform validate - Lint Kubernetes manifests:
helm lint - Follow existing naming conventions
- Document all variables and outputs
- Check Troubleshooting section
- Review Deployment Guide
- Consult Runbooks
- Check application/infrastructure logs
- Contact infrastructure team
For critical production issues, see Runbooks Escalation Procedures
- Prometheus and Grafana communities
- Google Cloud Platform documentation
- Terraform and Helm communities
- SRE best practices from Google SRE Book
| Version | Date | Changes |
|---|---|---|
| 1.0.0 | 2025-11-02 | Initial release with complete infrastructure and monitoring |
Maintained By: Infrastructure Team
Last Updated: 2025-11-02
Status: Production Ready β