ShivamMathtech/ML-Learning-Repository

The capstone models future churn from eight current customer snapshot features. It validates input, removes duplicate evidence, creates independent train/validation/test partitions, fits preprocessing within training folds, compares models, searches hyperparameters, selects a decision threshold on validation data, and evaluates the frozen procedure

★ 1Forks 0Jupyter NotebookGitHub ↗Compare

README

ML Learning Repository

Learn the mathematics. Build the workflow. Inspect the evidence.

A progressive machine-learning course and runnable customer-churn capstone in one Python repository. Includes 22 notebooks, 8 NumPy reference algorithms, 8 practice projects, a reusable training pipeline, FastAPI, a Streamlit dashboard, local experiment artifacts, optional MLflow, and tests.

The main experiment uses synthetic, generated customer snapshots. It is an engineering and learning reference, not a validated business decision system. Read VALIDATION.md for what was actually executed in the delivery environment, including optional tools and container limitations.

Held-out evaluation

Overview

The capstone models future churn from eight current customer snapshot features. It validates input, removes duplicate evidence, creates independent train/validation/test partitions, fits preprocessing within training folds, compares models, searches hyperparameters, selects a decision threshold on validation data, and evaluates the frozen procedure on the final test set. A Predictor loads the same preprocessing/model artifact for command-line, API and dashboard use.

Learning objectives

  • Read and manipulate data with Python, NumPy and Pandas.
  • Explain regression, classification, clustering, PCA and optimization with equations and code.
  • Design valid evaluation, recognize leakage, and use metrics appropriate to the task.
  • Build reusable pipelines and versioned model artifacts with strict input contracts.
  • Explore explanations, costs, calibration, time-series evaluation and drift.
  • Extend the code through optional gradient boosting, Optuna, SHAP, MLflow and PyTorch modules.

Architecture

flowchart TD
    A["Customer snapshots"] --> B["Validate and remove duplicates"]
    B --> C["Independent partitions"]
    C --> D["Train: fit pipeline and search in CV"]
    C --> E["Validation: select threshold"]
    C --> F["Test: frozen evaluation"]
    D --> E
    E --> F
    F --> G["Versioned artifact and reports"]
    G --> H["Predictor"]
    H --> I["API and dashboard"]
    G --> J["Reference profile"]
    J --> K["Batch drift monitoring"]
Loading

The saved pipeline is ChurnFeatures → ColumnTransformer → estimator. The final model remains fitted on the training partition; it is not refitted on validation data after choosing the threshold. That makes the selected threshold and model consistent. Retraining on more data would require a new calibration/threshold protocol. EDA and permutation importance use training and validation data, respectively. The final test set does not enter any fitting or selection function.

Repository structure

Directory What you will find
src/ml_project/data/ Offline generation, schema validation, cleanup, stratified splits
src/ml_project/features/ Snapshot features, imputers, scalers, encoders, selectors
src/ml_project/models/ Classification, regression, ensembles, clustering, reduction, optional PyTorch
src/ml_project/scratch/ Eight readable NumPy reference implementations
src/ml_project/training/ Metrics, comparison, tuning, CV, calibration and tracking
src/ml_project/pipelines/ Complete versioned customer-churn workflow
src/ml_project/inference/, api/ Predictor, typed requests, readiness, access key support
src/ml_project/monitoring/ Numeric and categorical batch-drift checks
app/ Streamlit exploration, prediction and explanation dashboard
notebooks/ 15 core lessons and 7 extended lessons
projects/ Eight practice-project guides using reusable runnable code
docs/ Mathematics, algorithm guides, operations, curriculum and data policy
tests/ Unit, pipeline, API and dashboard tests
reports/sample/ Executed capstone example: metrics, figures and prediction samples
configs/ Data, model, training and output configuration

Installation and environment setup

Use Python 3.11 or 3.12 for the quickest path. Python 3.13 support depends on optional wheels. Open a terminal inside the extracted ml-learning-repository directory.

Windows PowerShell:

py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[dev,notebooks,dashboard]"

If activation is restricted, use .\.venv\Scripts\python.exe instead of python in commands. You can run everything without changing your system execution policy.

Linux / macOS:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[dev,notebooks,dashboard]'

For the smallest working CLI/API environment: python -m pip install -e .. Optional extras are deliberately separate:

python -m pip install -e '.[advanced]'   # boosting, Optuna, SHAP, UMAP, resampling
python -m pip install -e '.[tracking]'   # MLflow tracking and UI
python -m pip install -e '.[deep]'       # PyTorch MLP/CNN

The first installation needs internet access to package indexes. The core dataset and experiments need no dataset download. Major dependency purposes are documented in docs/dependencies.md. The tested package versions are recorded in reports/environment_versions.json; version ranges in pyproject.toml are not an exact cross-platform lock. A tested direct-dependency constraints file is included as constraints-tested.txt. It does not pin every transitive dependency.

Dataset

Run ml-lab data to generate 2,400 deterministic rows with seed 42. Repeating it preserves an existing CSV; ml-lab data --force explicitly regenerates it. download_data.py is an offline generator, despite the conventional script name. The data contains realistic missingness and noisy outcomes, but no real people. See docs/data_card.md for schema, generator assumptions and limits.

You may set data.path to your own schema-compatible CSV. The loader does not silently recode categories or invent targets. Validate first with ml-lab validate. customer_id is metadata and is excluded from features. Prediction batches contain exactly the eight feature columns, without IDs or targets.

Quick start

ml-lab pipeline
ml-lab predict --input examples/predict_request.json --output reports/my_prediction.json
python -m pytest

The pipeline generates data if absent, trains dummy/logistic/random-forest/gradient-boosting candidates, searches the selected model, creates figures, saves reports, and publishes a local model version. Expect runtime to vary with CPU and package versions. ml-lab pipeline --quick uses smaller models, no tuning, and no new figures for a fast smoke run.

Training

python scripts/download_data.py
python scripts/validate_data.py
python scripts/train.py
# Equivalent:
ml-lab train --config configs/config.yaml

Edit configs/model.yaml to change candidates or parameters. Names include dummy, logistic, knn, naive_bayes, decision_tree, random_forest, gradient_boosting, svm, and optional xgboost, lightgbm, catboost. All requested external boosting models are implemented behind the advanced extra.

Edit configs/training.yaml for folds, seed, search method (grid, random, optuna), iterations, threshold objective (f1 or cost), costs and figure generation. Regression estimators and reusable comparison are demonstrated in notebook 07 and practice projects; the deployable API contract is specifically customer churn, not an unrestricted arbitrary-model API.

Each run writes reports/runs/<version>/ with data validation, split IDs, CSV partitions, configuration, comparison, search/threshold metadata, evaluation, importance, figures, and the complete model. models/production/<version>/ contains the serving files; current.json atomically selects the version. The API loads that version at startup. Restart it after training to adopt a new version.

Evaluation

ml-lab evaluate
ml-lab evaluate --input path/to/new_labeled_customers.csv --output reports/new_evaluation

Default evaluation reloads the current artifact and its recorded test split; it does not retrain. The binary report includes accuracy, precision, recall, F1, specificity, ROC-AUC, trapezoidal PR-AUC, average precision, log loss, threshold, sample count and a confusion matrix. Undefined scores use JSON null. Regression reports include MAE, MSE, RMSE, R², adjusted R² when applicable, and MAPE only when nonzero targets permit it. See docs/metrics.md.

Prediction

ml-lab predict --input examples/predict_request.json

The Predictor accepts a record, list of records, or a DataFrame. Fields may be explicitly null, but a record with all fields missing is rejected. Unknown fields, unknown categories, invalid ranges and nonfinite values fail with useful errors. Batches are limited to 1,000 records. Scores are estimated probabilities, not guarantees or calibrated confidence intervals.

from ml_project.inference.predictor import Predictor
from ml_project.utils.io import read_json

predictor = Predictor("models/production")
response = predictor.predict(read_json("examples/predict_request.json")["records"])

Only load trusted joblib artifacts. A checksum detects accidental corruption; it does not make pickle safe. A different scikit-learn version triggers a clear retraining instruction.

API

uvicorn ml_project.api.main:app --host 127.0.0.1 --port 8000

Open http://127.0.0.1:8000/docs for the interactive schema and requests.

Endpoint Behavior
GET / Service information
GET /health 200 when the model is loaded; 503 when unavailable
GET /model-info Version, feature contract, threshold and model identity
POST /predict Validated batch predictions with model version
curl -X POST http://127.0.0.1:8000/predict \
  -H 'Content-Type: application/json' \
  --data-binary @examples/predict_request.json

PowerShell users can use curl.exe with the same arguments. See examples/predict_response.json for an actual recorded response. Its score and version refer to the bundled example experiment, not future runs. Optional ML_API_KEY protects prediction and model-info endpoints via an X-API-Key header. The local tutorial defaults to no key. Configure ML_MODEL_DIR to select a different model directory.

Dashboard

streamlit run app/streamlit_app.py --browser.gatherUsageStats false

Open http://localhost:8501. Tabs cover training-data EDA, CV comparison, final evaluation, individual and CSV batch predictions, permutation importance, partial dependence/ICE, optional SHAP, and batch drift. The dashboard loads the same artifact; it never trains on a button press. Run a full pipeline if you want static ROC/PR figures after a quick run.

MLflow

python -m pip install -e '.[tracking]'
ml-lab train --tracking
mlflow ui --backend-store-uri sqlite:///mlflow.db --host 127.0.0.1 --port 5000

Tracking records parameters, metrics, model artifacts, configuration, dataset SHA-256, source commit, random seed and training duration. MLFLOW_TRACKING_URI overrides the YAML value. Tracking is off by default; local metadata is always saved. Enabling tracking without the dependency fails clearly. A source ZIP has no commit, so metadata says unversioned-source-archive until you initialize Git. MLflow tracking is implemented; an automated production registry approval workflow is not included.

Docker

docker compose up --build

The training service completes before API and dashboard start. Named volumes retain artifacts and reports. API and dashboard bind host ports to localhost. A non-root user runs the application. For a manual image workflow:

docker build -t ml-learning:local .
docker run --rm -p 127.0.0.1:8000:8000 -v "$PWD/models:/app/models:ro" ml-learning:local

Train locally before the manual inference-only command. Retrain with the container's dependencies if saved-model versions differ. The Docker build was not executed when Docker was unavailable in the delivery environment; see VALIDATION.md. Container files do not constitute an actual deployment.

Testing and quality

python -m pytest
python -m ruff check .
python -m ruff format --check .

Tests cover schema failures, leakage-resistant split behavior, imputer state, NumPy/sklearn comparisons, known metrics, saved-model predictions, corruption, no fitting at inference, readiness, API payloads, optional API keys, drift, and dashboard interactions. Ruff performs linting, import sorting and formatting, avoiding duplicate Black/isort dependencies. GitHub Actions runs tests and code checks. No cloud deployment or publishing is performed by CI.

Notebooks and practice projects

jupyter lab
python scripts/execute_notebooks.py
python scripts/run_projects.py all
python scripts/run_projects.py time_series

The standard executor uses a fresh Jupyter kernel for each lesson, fails on real errors, and records its results. A socket-free alternative, python scripts/execute_notebooks_inprocess.py, runs each notebook in a fresh IPython process and saves its actual outputs. This alternative was used for the bundled executed lessons because ZeroMQ kernel transport was blocked in the delivery environment. The optional deep notebooks explicitly report when PyTorch computation was not executed. If your kernel does not match the activated environment, register it:

python -m ipykernel install --user --name ml-learning --display-name 'ML Learning'
python scripts/execute_notebooks.py --kernel ml-learning

For deep learning:

python scripts/run_deep_learning.py --kind mlp --epochs 20
python scripts/run_deep_learning.py --kind cnn --epochs 20

The MLP and CNN use bundled 8×8 digit images, CPU training, DataLoader, validation loss, checkpointing and early stopping. See docs/deep_learning.md.

Monitoring

ml-lab monitor --input path/to/current_batch.csv --output reports/drift.json

Numeric PSI, KS p-values, missingness changes and categorical total variation compare a current batch to a stored training reference. Alert thresholds are illustrative. Feature drift does not prove performance decay; delayed labels are needed to evaluate actual error and calibration. The example logs request batch size, elapsed time and model version without logging customer payloads.

Common problems

Symptom Resolution
No module named ml_project Activate the intended environment and run python -m pip install -e . from the root
Missing configuration Run from the extracted project root or supply --config / ML_CONFIG
API returns 503 Train first, check ML_MODEL_DIR, inspect startup logs, restart after training
sklearn version mismatch Retrain with the serving environment; do not suppress the compatibility check
Optional import unavailable Install the appropriate extra; core ML does not require it
Kernel package missing Register the current environment with ipykernel
Port already used Choose another --port or Streamlit --server.port
Inference schema rejected Match the eight fields and declared categories in the example JSON
LightGBM native library error Install the platform OpenMP runtime or use the core sklearn models
Slow notebook run Start with core lessons and smaller sample counts; external models are optional

Learning roadmap / beginner → advanced path

Stage Lessons Build confidence by
Foundations 01–03, 17 Reading arrays, tables, equations and validation errors
Data and representation 04–06, 16 Answering EDA questions and proving leakage protection
Classical ML 07–11, 18 Comparing algorithms, diagnostics and NumPy implementations
Selection and explanation 12–14, 19 Tuning within CV, explaining scores and calibrating decisions
Capstone and operations 15, 20 Reproducing an experiment, serving it and monitoring inputs
Optional neural networks 21–22 Training MLP/CNN models with validation-based checkpoints

Each lesson includes explanations, executable examples, a visual where useful, exercises, common mistakes and interview questions. docs/learning_roadmap.md gives a study plan; docs/algorithms/ connects individual methods to code and extensions.

Future improvements

Extend the educational foundation with real prediction-horizon definitions, grouped or temporal customer splits, representative data, calibrated business costs, service-level monitoring, access management, load testing, delayed-label evaluation and reviewed model promotion. Those are explicit extensions, not claims about this sample repository's readiness for public production traffic.

License

MIT for original code, documentation and generated synthetic data. Third-party libraries retain their own licenses. Bundled sklearn datasets used in demonstrations retain their upstream provenance; docs/data_card.md explains where they are used.

Contributors

ShivamMathtech

Issues