GoGstickGo/transaction-api-deployment

Transaction api deployment

β˜… 0Forks 0HCLGitHub β†—Compare

README

Transaction API

Complete infrastructure-as-code setup for deploying a production-ready Transaction API on Google Kubernetes Engine (GKE) with comprehensive monitoring and observability.


πŸ—οΈ Overview

This repository contains Terraform modules and Helm charts for deploying:

  • GKE Cluster - Google Kubernetes Engine cluster with autoscaling
  • Transaction API - RESTful API for transaction processing
  • PostgreSQL Database - Persistent data storage
  • Monitoring Stack - Prometheus, Grafana, AlertManager
  • Observability - Metrics, alerts, dashboards, and SLO tracking

πŸ“ Repository Structure

.
β”œβ”€β”€ app/
β”‚   └── transaction-api/          # Transaction API application deployment
β”œβ”€β”€ infrastructure/
β”‚   β”œβ”€β”€ cluster/                  # GKE cluster infrastructure
β”‚   β”œβ”€β”€ image-repo/               # Docker image artifact registry
β”‚   └── monitoring/               # Prometheus monitoring stack
β”œβ”€β”€ modules/
β”‚   β”œβ”€β”€ cluster/                  # Reusable GKE cluster module
β”‚   β”œβ”€β”€ image-repo/               # Reusable image repository module
β”‚   β”œβ”€β”€ postgresql/               # PostgreSQL Helm chart module
β”‚   β”œβ”€β”€ prometheus/               # Prometheus stack module
β”‚   └── transaction-api/          # Transaction API Helm chart module
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ Deployment/               # Deployment guides
β”‚   β”œβ”€β”€ Monitoring/               # Monitoring setup and guides
β”‚   └── RunBooks/                 # Operational runbooks
└── README.md                     # This file

πŸ“š Documentation

πŸš€ Deployment

Document Description
DEPLOYMENT.md Complete step-by-step deployment guide with prerequisites, deployment order, verification steps, and troubleshooting

Key Topics:

  • Prerequisites and tool setup
  • Infrastructure deployment (image-repo, cluster, monitoring)
  • Application deployment (Transaction API + PostgreSQL)
  • Verification procedures
  • Rollback strategies
  • Cleanup/teardown instructions

πŸ“Š Monitoring

Document Description
README.md Monitoring overview, architecture, and quick start guide
transaction-api-monitoring-guide.md Complete monitoring implementation guide with SLOs, alerts, and dashboards
QUICK_REFERENCE.md Quick reference card for SLOs, metrics queries, alerts, and troubleshooting

Key Topics:

  • Prometheus metrics collection
  • Grafana dashboard setup
  • Service Level Objectives (SLOs)
  • Alert rules and thresholds
  • Error budget tracking
  • Application instrumentation examples (Python/Node.js)

πŸ”§ Operational Runbooks

Runbook Alert Description
README.md - Runbooks overview and quick reference
high-error-rate.md TransactionAPIHighErrorRate Error rate > 1% - deployment issues, database problems, resource exhaustion
service-down.md TransactionAPIDown Complete service outage - pod crashes, scaling issues, network problems
database-errors.md DatabaseConnectionFailures DB error rate > 0.1% - connection pool, slow queries, network issues
high-latency.md TransactionAPIHighLatency P95 > 200ms - CPU pressure, slow queries, external API delays

Each runbook includes:

  • Alert description and thresholds
  • Quick diagnosis steps with PromQL queries
  • Common causes and solutions
  • Step-by-step investigation procedures
  • Verification steps
  • Escalation procedures

πŸ›οΈ Architecture

Infrastructure Components

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                      GCP Project                         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                           β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚   Artifact   β”‚  β”‚  GCS Bucket  β”‚  β”‚ GKE Cluster  β”‚  β”‚
β”‚  β”‚   Registry   β”‚  β”‚  (TF State)  β”‚  β”‚              β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                               β”‚           β”‚
β”‚       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜          β”‚
β”‚       β”‚                                                   β”‚
β”‚  β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚           GKE Cluster Namespaces                   β”‚  β”‚
β”‚  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”              β”‚  β”‚
β”‚  β”‚  β”‚transactions  β”‚  β”‚  monitoring   β”‚              β”‚  β”‚
β”‚  β”‚  β”‚              β”‚  β”‚               β”‚              β”‚  β”‚
β”‚  β”‚  β”‚β€’ Trans API   β”‚  β”‚β€’ Prometheus   β”‚              β”‚  β”‚
β”‚  β”‚  β”‚β€’ PostgreSQL  β”‚  β”‚β€’ Grafana      β”‚              β”‚  β”‚
β”‚  β”‚  β”‚              β”‚  β”‚β€’ AlertManager β”‚              β”‚  β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Monitoring Flow

Transaction API (App)
         ↓ (exposes /metrics)
    Prometheus (scrapes metrics)
         ↓
    β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”
    ↓         ↓
Grafana   AlertManager
(visualize) (notify)

πŸ”‘ Key Features

Infrastructure

βœ… Multi-zone GKE cluster with autoscaling
βœ… Managed node pools with auto-repair and auto-upgrade
βœ… Artifact Registry for Docker images
βœ… Infrastructure as Code - Complete Terraform setup
βœ… Makefile automation - Simple deployment commands

Application

βœ… Horizontal Pod Autoscaling (HPA)
βœ… Pod Disruption Budgets (PDB)
βœ… Health checks and liveness probes
βœ… Resource limits and requests
βœ… Pod anti-affinity for high availability


πŸ› οΈ Technologies Used

Category Technology
Cloud Provider Google Cloud Platform (GCP)
Container Orchestration Google Kubernetes Engine (GKE)
Infrastructure as Code Terraform
Package Management Helm
Monitoring Prometheus, Grafana, AlertManager
Database PostgreSQL
Programming Go (Transaction API)

🀝 Contributing

Development Workflow

  1. Create feature branch
  2. Make changes
  3. Test locally
  4. Run terraform plan to preview changes
  5. Submit pull request
  6. Deploy to staging first
  7. Verify in staging
  8. Deploy to production

Code Standards

  • Use Terraform formatting: terraform fmt -recursive
  • Validate Terraform: terraform validate
  • Lint Kubernetes manifests: helm lint
  • Follow existing naming conventions
  • Document all variables and outputs

πŸ“ž Support

Getting Help

  1. Check Troubleshooting section
  2. Review Deployment Guide
  3. Consult Runbooks
  4. Check application/infrastructure logs
  5. Contact infrastructure team

Escalation

For critical production issues, see Runbooks Escalation Procedures

πŸ™ Acknowledgments

  • Prometheus and Grafana communities
  • Google Cloud Platform documentation
  • Terraform and Helm communities
  • SRE best practices from Google SRE Book

πŸ“… Version History

Version Date Changes
1.0.0 2025-11-02 Initial release with complete infrastructure and monitoring

Maintained By: Infrastructure Team
Last Updated: 2025-11-02
Status: Production Ready βœ…

Contributors

GoGstickGo

Issues