hard02/data-pipeline-cli

Enterprise Data Engineering, Quality, Governance & Lineage Platform built in Python.

β˜… 0Forks 0PythonGitHub β†—Compare

README

DataPipeline CLI πŸš€

Python Version Coverage License: MIT

DataPipeline CLI is an enterprise-grade Data Engineering, Quality, Governance, and Lineage platform. It provides a robust, strictly-typed Python foundation for automating the ingestion, validation, transformation, and auditing of data assets.

Built specifically for modern data platforms, it unifies the entire data lifecycle into a single, cohesive CLI and API ecosystem.


🌟 Key Features & Engines

This project is not just a script; it is a modular monolith composed of several highly specialized engines:

  • πŸ“₯ Ingestion Engine: Chunked, memory-efficient data loading with automatic schema inference.
  • πŸ›‘οΈ Validation Engine: Strict Pydantic-powered schema enforcement, referential integrity checks, and data quality profiling.
  • πŸ”„ Transformation Engine: Idempotent data transformations, type coercion, and atomic rollbacks.
  • πŸ•΅οΈ Anomaly Detection: Statistical analysis to identify data drift and quality outliers.
  • πŸ›οΈ Governance Engine: Master data management, stewardship tracking, and metadata tagging (GDPR/CCPA readiness).
  • πŸ•ΈοΈ Lineage Engine: Graph-based tracking of data provenance and transformation history.
  • πŸ”’ Audit Engine: Immutable, SHA-256 hashed event logs for regulatory compliance.
  • πŸ“Š Reporting Engine: Automated HTML, Markdown, and CSV generation for data quality reports.

πŸ—οΈ Architecture Stack

  • CLI Framework: Click & Rich for beautiful terminal UI.
  • API Layer: FastAPI for exposing core engines via REST.
  • Data Processing: Pandas for vectorized transformations.
  • Data Validation: Pydantic for runtime type checking.
  • Persistence: SQLite (with SQLAlchemy abstractions for future PostgreSQL scaling).

πŸš€ Quick Start

Installation

Ensure you have Python 3.10+ installed.

# Clone the repository
git clone https://github.com/hard02/data-pipeline-cli.git
cd data-pipeline-cli

# Create a virtual environment and install dependencies
python3 -m venv venv
source venv/bin/activate
pip install -e .[dev,stats]

Basic CLI Usage

The CLI exposes all core engines via intuitive commands.

# Ingest data
datapipeline ingest ./data/source.csv

# Validate data against strict rules
datapipeline validate ./data/source.csv

# Run the full orchestrator pipeline (Ingest -> Validate -> Transform -> Audit)
datapipeline pipeline-run ./data/source.csv --config ./configs/pipeline.yaml

πŸ§ͺ Testing and Quality

This project is built with test-driven principles, achieving 90.23% code coverage across 177 unit and integration tests.

# Run the test suite
make test
# OR
pytest tests/ -v --cov=src

πŸ“‚ Project Structure

data-pipeline-cli/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ api/            # FastAPI REST endpoints
β”‚   β”œβ”€β”€ cli/            # Click CLI commands and Display logic
β”‚   β”œβ”€β”€ pipeline/       # Core Orchestrator
β”‚   β”œβ”€β”€ ingestion/      # Data loading logic
β”‚   β”œβ”€β”€ validation/     # Pydantic schema rules
β”‚   β”œβ”€β”€ transformation/ # Data manipulation
β”‚   β”œβ”€β”€ anomaly/        # Statistical outlier detection
β”‚   β”œβ”€β”€ lineage/        # Graph-based provenance
β”‚   β”œβ”€β”€ governance/     # Stewardship tracking
β”‚   β”œβ”€β”€ audit/          # Immutable SHA-256 event logging
β”‚   β”œβ”€β”€ reporting/      # Document generation
β”‚   └── persistence/    # SQLite/SQLAlchemy models
β”œβ”€β”€ tests/              # 177+ Pytest unit & integration tests
β”œβ”€β”€ docs/               # Architecture and functional documentation
β”œβ”€β”€ configs/            # YAML pipeline configurations
└── pyproject.toml      # Project metadata and dependencies

πŸ“œ License

This project is licensed under the MIT License.

Contributors

hard02

Issues