A progressive machine-learning course and runnable customer-churn capstone in one Python repository. Includes 22 notebooks, 8 NumPy reference algorithms, 8 practice projects, a reusable training pipeline, FastAPI, a Streamlit dashboard, local experiment artifacts, optional MLflow, and tests.
The main experiment uses synthetic, generated customer snapshots. It is an engineering and learning reference, not a validated business decision system. Read VALIDATION.md for what was actually executed in the delivery environment, including optional tools and container limitations.
The capstone models future churn from eight current customer snapshot features. It validates input, removes duplicate evidence, creates independent train/validation/test partitions, fits preprocessing within training folds, compares models, searches hyperparameters, selects a decision threshold on validation data, and evaluates the frozen procedure on the final test set. A Predictor loads the same preprocessing/model artifact for command-line, API and dashboard use.
- Read and manipulate data with Python, NumPy and Pandas.
- Explain regression, classification, clustering, PCA and optimization with equations and code.
- Design valid evaluation, recognize leakage, and use metrics appropriate to the task.
- Build reusable pipelines and versioned model artifacts with strict input contracts.
- Explore explanations, costs, calibration, time-series evaluation and drift.
- Extend the code through optional gradient boosting, Optuna, SHAP, MLflow and PyTorch modules.
flowchart TD
A["Customer snapshots"] --> B["Validate and remove duplicates"]
B --> C["Independent partitions"]
C --> D["Train: fit pipeline and search in CV"]
C --> E["Validation: select threshold"]
C --> F["Test: frozen evaluation"]
D --> E
E --> F
F --> G["Versioned artifact and reports"]
G --> H["Predictor"]
H --> I["API and dashboard"]
G --> J["Reference profile"]
J --> K["Batch drift monitoring"]
The saved pipeline is ChurnFeatures → ColumnTransformer → estimator. The final model remains fitted
on the training partition; it is not refitted on validation data after choosing the threshold.
That makes the selected threshold and model consistent. Retraining on more data would require a
new calibration/threshold protocol. EDA and permutation importance use training and validation data,
respectively. The final test set does not enter any fitting or selection function.
| Directory | What you will find |
|---|---|
src/ml_project/data/ |
Offline generation, schema validation, cleanup, stratified splits |
src/ml_project/features/ |
Snapshot features, imputers, scalers, encoders, selectors |
src/ml_project/models/ |
Classification, regression, ensembles, clustering, reduction, optional PyTorch |
src/ml_project/scratch/ |
Eight readable NumPy reference implementations |
src/ml_project/training/ |
Metrics, comparison, tuning, CV, calibration and tracking |
src/ml_project/pipelines/ |
Complete versioned customer-churn workflow |
src/ml_project/inference/, api/ |
Predictor, typed requests, readiness, access key support |
src/ml_project/monitoring/ |
Numeric and categorical batch-drift checks |
app/ |
Streamlit exploration, prediction and explanation dashboard |
notebooks/ |
15 core lessons and 7 extended lessons |
projects/ |
Eight practice-project guides using reusable runnable code |
docs/ |
Mathematics, algorithm guides, operations, curriculum and data policy |
tests/ |
Unit, pipeline, API and dashboard tests |
reports/sample/ |
Executed capstone example: metrics, figures and prediction samples |
configs/ |
Data, model, training and output configuration |
Use Python 3.11 or 3.12 for the quickest path. Python 3.13 support depends on optional wheels.
Open a terminal inside the extracted ml-learning-repository directory.
Windows PowerShell:
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[dev,notebooks,dashboard]"If activation is restricted, use .\.venv\Scripts\python.exe instead of python in commands.
You can run everything without changing your system execution policy.
Linux / macOS:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[dev,notebooks,dashboard]'For the smallest working CLI/API environment: python -m pip install -e ..
Optional extras are deliberately separate:
python -m pip install -e '.[advanced]' # boosting, Optuna, SHAP, UMAP, resampling
python -m pip install -e '.[tracking]' # MLflow tracking and UI
python -m pip install -e '.[deep]' # PyTorch MLP/CNNThe first installation needs internet access to package indexes. The core dataset and experiments
need no dataset download. Major dependency purposes are documented in docs/dependencies.md.
The tested package versions are recorded in reports/environment_versions.json; version ranges in
pyproject.toml are not an exact cross-platform lock. A tested direct-dependency constraints file is
included as constraints-tested.txt. It does not pin every transitive dependency.
Run ml-lab data to generate 2,400 deterministic rows with seed 42. Repeating it preserves an existing
CSV; ml-lab data --force explicitly regenerates it. download_data.py is an offline generator, despite
the conventional script name. The data contains realistic missingness and noisy outcomes, but no
real people. See docs/data_card.md for schema, generator assumptions and limits.
You may set data.path to your own schema-compatible CSV. The loader does not silently recode categories
or invent targets. Validate first with ml-lab validate. customer_id is metadata and is excluded from
features. Prediction batches contain exactly the eight feature columns, without IDs or targets.
ml-lab pipeline
ml-lab predict --input examples/predict_request.json --output reports/my_prediction.json
python -m pytestThe pipeline generates data if absent, trains dummy/logistic/random-forest/gradient-boosting candidates,
searches the selected model, creates figures, saves reports, and publishes a local model version.
Expect runtime to vary with CPU and package versions. ml-lab pipeline --quick uses smaller models,
no tuning, and no new figures for a fast smoke run.
python scripts/download_data.py
python scripts/validate_data.py
python scripts/train.py
# Equivalent:
ml-lab train --config configs/config.yamlEdit configs/model.yaml to change candidates or parameters. Names include dummy, logistic, knn,
naive_bayes, decision_tree, random_forest, gradient_boosting, svm, and optional xgboost, lightgbm,
catboost. All requested external boosting models are implemented behind the advanced extra.
Edit configs/training.yaml for folds, seed, search method (grid, random, optuna), iterations,
threshold objective (f1 or cost), costs and figure generation. Regression estimators and reusable
comparison are demonstrated in notebook 07 and practice projects; the deployable API contract is
specifically customer churn, not an unrestricted arbitrary-model API.
Each run writes reports/runs/<version>/ with data validation, split IDs, CSV partitions, configuration,
comparison, search/threshold metadata, evaluation, importance, figures, and the complete model.
models/production/<version>/ contains the serving files; current.json atomically selects the version.
The API loads that version at startup. Restart it after training to adopt a new version.
ml-lab evaluate
ml-lab evaluate --input path/to/new_labeled_customers.csv --output reports/new_evaluationDefault evaluation reloads the current artifact and its recorded test split; it does not retrain.
The binary report includes accuracy, precision, recall, F1, specificity, ROC-AUC, trapezoidal PR-AUC,
average precision, log loss, threshold, sample count and a confusion matrix. Undefined scores use JSON
null. Regression reports include MAE, MSE, RMSE, R², adjusted R² when applicable, and MAPE only when
nonzero targets permit it. See docs/metrics.md.
ml-lab predict --input examples/predict_request.jsonThe Predictor accepts a record, list of records, or a DataFrame. Fields may be explicitly null, but a record with all fields missing is rejected. Unknown fields, unknown categories, invalid ranges and nonfinite values fail with useful errors. Batches are limited to 1,000 records. Scores are estimated probabilities, not guarantees or calibrated confidence intervals.
from ml_project.inference.predictor import Predictor
from ml_project.utils.io import read_json
predictor = Predictor("models/production")
response = predictor.predict(read_json("examples/predict_request.json")["records"])Only load trusted joblib artifacts. A checksum detects accidental corruption; it does not make pickle safe. A different scikit-learn version triggers a clear retraining instruction.
uvicorn ml_project.api.main:app --host 127.0.0.1 --port 8000Open http://127.0.0.1:8000/docs for the interactive schema and requests.
| Endpoint | Behavior |
|---|---|
GET / |
Service information |
GET /health |
200 when the model is loaded; 503 when unavailable |
GET /model-info |
Version, feature contract, threshold and model identity |
POST /predict |
Validated batch predictions with model version |
curl -X POST http://127.0.0.1:8000/predict \
-H 'Content-Type: application/json' \
--data-binary @examples/predict_request.jsonPowerShell users can use curl.exe with the same arguments. See examples/predict_response.json for
an actual recorded response. Its score and version refer to the bundled example experiment, not future
runs. Optional ML_API_KEY protects prediction and model-info endpoints via an X-API-Key header.
The local tutorial defaults to no key. Configure ML_MODEL_DIR to select a different model directory.
streamlit run app/streamlit_app.py --browser.gatherUsageStats falseOpen http://localhost:8501. Tabs cover training-data EDA, CV comparison, final evaluation, individual
and CSV batch predictions, permutation importance, partial dependence/ICE, optional SHAP, and batch
drift. The dashboard loads the same artifact; it never trains on a button press. Run a full pipeline
if you want static ROC/PR figures after a quick run.
python -m pip install -e '.[tracking]'
ml-lab train --tracking
mlflow ui --backend-store-uri sqlite:///mlflow.db --host 127.0.0.1 --port 5000Tracking records parameters, metrics, model artifacts, configuration, dataset SHA-256, source commit,
random seed and training duration. MLFLOW_TRACKING_URI overrides the YAML value. Tracking is off by
default; local metadata is always saved. Enabling tracking without the dependency fails clearly.
A source ZIP has no commit, so metadata says unversioned-source-archive until you initialize Git.
MLflow tracking is implemented; an automated production registry approval workflow is not included.
docker compose up --buildThe training service completes before API and dashboard start. Named volumes retain artifacts and reports. API and dashboard bind host ports to localhost. A non-root user runs the application. For a manual image workflow:
docker build -t ml-learning:local .
docker run --rm -p 127.0.0.1:8000:8000 -v "$PWD/models:/app/models:ro" ml-learning:localTrain locally before the manual inference-only command. Retrain with the container's dependencies if saved-model versions differ. The Docker build was not executed when Docker was unavailable in the delivery environment; see VALIDATION.md. Container files do not constitute an actual deployment.
python -m pytest
python -m ruff check .
python -m ruff format --check .Tests cover schema failures, leakage-resistant split behavior, imputer state, NumPy/sklearn comparisons, known metrics, saved-model predictions, corruption, no fitting at inference, readiness, API payloads, optional API keys, drift, and dashboard interactions. Ruff performs linting, import sorting and formatting, avoiding duplicate Black/isort dependencies. GitHub Actions runs tests and code checks. No cloud deployment or publishing is performed by CI.
jupyter lab
python scripts/execute_notebooks.py
python scripts/run_projects.py all
python scripts/run_projects.py time_seriesThe standard executor uses a fresh Jupyter kernel for each lesson, fails on real errors, and records its results.
A socket-free alternative, python scripts/execute_notebooks_inprocess.py, runs each notebook in a fresh
IPython process and saves its actual outputs. This alternative was used for the bundled executed lessons
because ZeroMQ kernel transport was blocked in the delivery environment.
The optional deep notebooks explicitly report when PyTorch computation was not executed. If your
kernel does not match the activated environment, register it:
python -m ipykernel install --user --name ml-learning --display-name 'ML Learning'
python scripts/execute_notebooks.py --kernel ml-learningFor deep learning:
python scripts/run_deep_learning.py --kind mlp --epochs 20
python scripts/run_deep_learning.py --kind cnn --epochs 20The MLP and CNN use bundled 8×8 digit images, CPU training, DataLoader, validation loss, checkpointing and early stopping. See docs/deep_learning.md.
ml-lab monitor --input path/to/current_batch.csv --output reports/drift.jsonNumeric PSI, KS p-values, missingness changes and categorical total variation compare a current batch to a stored training reference. Alert thresholds are illustrative. Feature drift does not prove performance decay; delayed labels are needed to evaluate actual error and calibration. The example logs request batch size, elapsed time and model version without logging customer payloads.
| Symptom | Resolution |
|---|---|
No module named ml_project |
Activate the intended environment and run python -m pip install -e . from the root |
| Missing configuration | Run from the extracted project root or supply --config / ML_CONFIG |
| API returns 503 | Train first, check ML_MODEL_DIR, inspect startup logs, restart after training |
| sklearn version mismatch | Retrain with the serving environment; do not suppress the compatibility check |
| Optional import unavailable | Install the appropriate extra; core ML does not require it |
| Kernel package missing | Register the current environment with ipykernel |
| Port already used | Choose another --port or Streamlit --server.port |
| Inference schema rejected | Match the eight fields and declared categories in the example JSON |
| LightGBM native library error | Install the platform OpenMP runtime or use the core sklearn models |
| Slow notebook run | Start with core lessons and smaller sample counts; external models are optional |
| Stage | Lessons | Build confidence by |
|---|---|---|
| Foundations | 01–03, 17 | Reading arrays, tables, equations and validation errors |
| Data and representation | 04–06, 16 | Answering EDA questions and proving leakage protection |
| Classical ML | 07–11, 18 | Comparing algorithms, diagnostics and NumPy implementations |
| Selection and explanation | 12–14, 19 | Tuning within CV, explaining scores and calibrating decisions |
| Capstone and operations | 15, 20 | Reproducing an experiment, serving it and monitoring inputs |
| Optional neural networks | 21–22 | Training MLP/CNN models with validation-based checkpoints |
Each lesson includes explanations, executable examples, a visual where useful, exercises, common mistakes and interview questions. docs/learning_roadmap.md gives a study plan; docs/algorithms/ connects individual methods to code and extensions.
Extend the educational foundation with real prediction-horizon definitions, grouped or temporal customer splits, representative data, calibrated business costs, service-level monitoring, access management, load testing, delayed-label evaluation and reviewed model promotion. Those are explicit extensions, not claims about this sample repository's readiness for public production traffic.
MIT for original code, documentation and generated synthetic data. Third-party libraries retain their own licenses. Bundled sklearn datasets used in demonstrations retain their upstream provenance; docs/data_card.md explains where they are used.
