shmWorks/Demand_Forecasting_System

★ 0Forks 0PythonGitHub ↗Compare

README

Retail-IQ: Multi-Family Demand Forecasting System

Python 3.10 Status License: MIT

Overview

Retail chains face a dual failure mode: overstocking perishable goods creates waste, while understocking drives revenue leakage. Existing tools often apply univariate models that ignore critical retail dynamics.

Retail-IQ is a multi-family panel regression forecasting system built to tackle this challenge. It analyzes daily sales across 54 stores and 33 product families using the Corporación Favorita Store Sales dataset from Kaggle. Beyond mere forecasting, the system quantifies promotional lift and detects cannibalization across adjacent SKUs.

Key Capabilities

  • Robust Pipeline: Highly optimized feature engineering using vectorized operations (O(N log K) time complexity for complex mappings).
  • Dual Tracking: Models continuous volume while evaluating promotional efficacy.
  • From-Scratch Engineering: Custom JAX-accelerated Gradient Descent Linear Regression baseline alongside standard advanced methods.
  • Scalable State: Designed using immutable, pure-function pipelines tracking strict zero-retention invariant strategies.

System Architecture

Retail-IQ is designed as a strict Directed Acyclic Graph (DAG) pipeline. Every stage is a pure function.

graph TD;
    A[Raw CSVs] -->|config.py paths| B(preprocessing.py: Parquet loader);
    B -->|ffill / merge| C(features.py: FastFeatureEngineer);
    C -->|"Vectorized O(N log K) Ops"| D(models.py: GD_Linear, SeasonalNaive);
    C -->|Optuna Tuned| E(XGBoost / LightGBM Notebooks);
    D --> F(evaluation.py: RMSLE, SHAP);
    E --> F;
Loading

See src/retail_iq/config.py for data path strictness and Learn_minmax_insights/01_System_Architecture.md for architectural invariants.


⚡ Quick Start Tutorial

To execute the forecasting pipeline using our strictly encapsulated modules, follow these steps.

1. Environment Setup

We highly recommend using uv for lightning-fast package management.

uv venv
source .venv/bin/activate
uv pip install -e .

2. Loading & Preprocessing Data

Data loading heavily favors Parquet files for significant speed boosts.

from retail_iq.preprocessing import load_raw_data, preprocess_dates, merge_datasets

# 1. Load data
train, test, stores, oil, holidays, transactions = load_raw_data()

# 2. Date conversion
train, test, oil, holidays, transactions = preprocess_dates([train, test, oil, holidays, transactions])

# 3. Merge panel
df_merged = merge_datasets(train, stores, oil, holidays, transactions)

Explore preprocessing.py

3. Feature Engineering

Use the fluent FastFeatureEngineer API. Rule of thumb: Never call .copy() inside chained methods, and data is sorted purely once during initialization to avoid leakages.

from retail_iq.features import FastFeatureEngineer

ffe = FastFeatureEngineer(df_merged, transactions, oil, holidays, stores)

df_features = (ffe
    .add_temporal_features()
    .add_lag_and_rolling(lags=[1, 7, 14, 365], windows=[7, 14, 28])
    .add_onpromotion_features()
    .add_macroeconomic_features()
    .transform()
)

Explore features.py

4. Baseline Modeling (From Scratch)

Run our custom JAX-compiled Gradient Descent Linear Regressor.

from retail_iq.models import GD_Linear
import numpy as np

# Apply target transformation
y_log = np.log1p(y_train)

# Initialize and fit
model = GD_Linear(lr=0.001, iterations=1000, l2=0.01)
model.fit(X_train, y_log)

# Predict and inverse transform
preds = np.expm1(model.predict(X_test))

Explore models.py

📊 Retail-IQ Dashboard

The project includes a lightweight, real-time dashboard for monitoring sales trends and model performance.

1. Prerequisites

  • MongoDB: Ensure MongoDB is running locally on mongodb://localhost:27017/.

2. Seeding the Database

Before running the dashboard for the first time, populate it with mock/pipeline data:

python -m dashboard.seed_db

Note: This creates an admin user with credentials admin / admin.

3. Launching the Application

Run the Flask server from the project root:

python -m dashboard.app

The dashboard will be available at http://127.0.0.1:5000.


🚀 Full Stack Dashboard Setup Guide

Retail-IQ now ships with a brutally minimalistic, high-performance web dashboard (Flask + MongoDB + Alpine.js + Chart.js) to monitor the data science pipeline execution and visualize forecasting metrics in real-time.

1. Prerequisites

2. Environment & Dependencies

Activate your virtual environment and install the required dependencies (including the new web stack):

# Using uv (recommended)
uv venv
source .venv/bin/activate
uv pip install -e .

Required packages include: flask, flask-cors, pymongo, pydantic, PyJWT, bcrypt.

3. Database Setup (MongoDB)

Ensure MongoDB is running on your system. If you are using Linux:

sudo systemctl start mongod

To populate the database with mock metrics for the dashboard, run the seeding script:

PYTHONPATH=. python3 dashboard/seed_db.py

This will create the necessary collections, targeted indexes, and an initial admin user.

4. Downloading Raw Data

The pipeline requires the Corporación Favorita dataset to execute. Download it directly from Kaggle into the data/raw/ directory:

mkdir -p data/raw
kaggle competitions download -c store-sales-time-series-forecasting -p data/raw
unzip data/raw/store-sales-time-series-forecasting.zip -d data/raw
rm data/raw/store-sales-time-series-forecasting.zip

5. Running the Application

Start the Flask backend server:

PYTHONPATH=. python3 dashboard/app.py
  • Navigate to http://127.0.0.1:5000 in your web browser.
  • Log in using the default credentials (Username: admin / Password: admin if generated via seed_db.py).
  • Use the Pipeline Execution Engine panel to trigger the data processing, feature engineering, and EDA pipeline. Real-time artifacts will appear directly on the dashboard once completed.

Project Structure

Retail-IQ/
├── data/
│   ├── raw/                # Unmodified input CSVs
│   └── processed/          # Cleaned and featured datasets (Parquet)
├── docs/                   # Full academic and technical documentation
├── notebooks/              # Strictly driver-notebooks for execution
├── src/
│   └── retail_iq/
│       ├── config.py       # Centralized strict path constants
│       ├── preprocessing.py# Pure functional data cleaners
│       ├── features.py     # FastFeatureEngineer fluent pipeline
│       ├── models.py       # Custom baselines (GD_Linear, Naive)
│       ├── evaluation.py   # Validation and SHAP utility suite
│       └── perf_utils.py   # Profiling utilities
└── pyproject.toml          # Project configuration

Implementation Phases

  • Phase 1: Project setup & repository init
  • Phase 2: Data Ingestion & Validation (Preprocessing)
  • Phase 3: Exploratory Data Analysis (EDA)
  • Phase 4: Vectorized Feature Engineering
  • Phase 5: Baseline Modeling (GD Linear / Seasonal Naive)
  • Phase 6: Advanced Modeling + Optuna Tuning
  • Phase 7: Evaluation + SHAP Analysis
  • Phase 8: Cannibalization & Lift Analysis
  • Phase 9: Finalization & Report Generation (Current Target)

🚨 Known Issues & Future Directions

For collaborators entering the repository, there are critical architectural and logical blindspots to address:

Critical Known Issues

  1. Memory Blowup: Current add_lag_and_rolling() operations create intermediate dataframe state replications. Investigating zero-copy methodologies is paramount.
  2. Lag Feature Leakage Risk: Any un-shifted rolling operations leak the current period. Strict guardrails must be preserved within the FastFeatureEngineer pipeline to prevent accidental leakage.
  3. Oil Price Spurious Correlation: Because the oil market closes on weekends, ffill pushes Friday prices to Saturday/Sunday. Given oil is a macro proxy, this introduces a lagged correlation risk.

Future Optimization Directions

  • Two-Stage Zero-Inflation Modeling: Approximately 40-60% of our rows contain structural zero-sales. All current models (GD, XGB, LGBM) output continuous probabilities and systematically under-predict zero periods. Action item: Transition to a two-stage approach: A classifier predicting P(sales=0) followed by a regressor predicting E(sales|sales>0).
  • Dtype Bloat: Memory compression pipeline needs to aggressively transition generic float64/int64 matrices into float32/int16 categories upon loading.

Core Team

  • Ayesha Khalid (23L-0667)
  • Uma E Rubab (23L-0928)
  • Sheraz Malik (23L-0572)

Contributors

shmWorks

Issues