DataPipeline CLI is an enterprise-grade Data Engineering, Quality, Governance, and Lineage platform. It provides a robust, strictly-typed Python foundation for automating the ingestion, validation, transformation, and auditing of data assets.
Built specifically for modern data platforms, it unifies the entire data lifecycle into a single, cohesive CLI and API ecosystem.
This project is not just a script; it is a modular monolith composed of several highly specialized engines:
- π₯ Ingestion Engine: Chunked, memory-efficient data loading with automatic schema inference.
- π‘οΈ Validation Engine: Strict Pydantic-powered schema enforcement, referential integrity checks, and data quality profiling.
- π Transformation Engine: Idempotent data transformations, type coercion, and atomic rollbacks.
- π΅οΈ Anomaly Detection: Statistical analysis to identify data drift and quality outliers.
- ποΈ Governance Engine: Master data management, stewardship tracking, and metadata tagging (GDPR/CCPA readiness).
- πΈοΈ Lineage Engine: Graph-based tracking of data provenance and transformation history.
- π Audit Engine: Immutable, SHA-256 hashed event logs for regulatory compliance.
- π Reporting Engine: Automated HTML, Markdown, and CSV generation for data quality reports.
- CLI Framework: Click & Rich for beautiful terminal UI.
- API Layer: FastAPI for exposing core engines via REST.
- Data Processing: Pandas for vectorized transformations.
- Data Validation: Pydantic for runtime type checking.
- Persistence: SQLite (with SQLAlchemy abstractions for future PostgreSQL scaling).
Ensure you have Python 3.10+ installed.
# Clone the repository
git clone https://github.com/hard02/data-pipeline-cli.git
cd data-pipeline-cli
# Create a virtual environment and install dependencies
python3 -m venv venv
source venv/bin/activate
pip install -e .[dev,stats]The CLI exposes all core engines via intuitive commands.
# Ingest data
datapipeline ingest ./data/source.csv
# Validate data against strict rules
datapipeline validate ./data/source.csv
# Run the full orchestrator pipeline (Ingest -> Validate -> Transform -> Audit)
datapipeline pipeline-run ./data/source.csv --config ./configs/pipeline.yamlThis project is built with test-driven principles, achieving 90.23% code coverage across 177 unit and integration tests.
# Run the test suite
make test
# OR
pytest tests/ -v --cov=srcdata-pipeline-cli/
βββ src/
β βββ api/ # FastAPI REST endpoints
β βββ cli/ # Click CLI commands and Display logic
β βββ pipeline/ # Core Orchestrator
β βββ ingestion/ # Data loading logic
β βββ validation/ # Pydantic schema rules
β βββ transformation/ # Data manipulation
β βββ anomaly/ # Statistical outlier detection
β βββ lineage/ # Graph-based provenance
β βββ governance/ # Stewardship tracking
β βββ audit/ # Immutable SHA-256 event logging
β βββ reporting/ # Document generation
β βββ persistence/ # SQLite/SQLAlchemy models
βββ tests/ # 177+ Pytest unit & integration tests
βββ docs/ # Architecture and functional documentation
βββ configs/ # YAML pipeline configurations
βββ pyproject.toml # Project metadata and dependencies
This project is licensed under the MIT License.