Repository for the recommander systems project source code
gantt
title Outfit Recommendation App - 4 Week Sprint
dateFormat YYYY-MM-DD
axisFormat %d-%b
section Week 1: Blueprint & Pitch
Team Assignment & Roles :done, 2025-11-04, 1d
Define Problem & Target Users :done, 2025-11-04, 2d
Choose Dataset & Data Strategy :done, 2025-11-05, 3d
Model Innovation Selection :done, 2025-11-07, 2d
Technical Stack & Wireframes :done, 2025-11-09, 2d
Project Plan PDF (Deliverable) :milestone, 2025-11-11, 1d
section Week 2: First Build
Backend Skeleton (FastAPI) :active, 2025-11-12, 3d
Frontend MVP (Streamlit) :active, 2025-11-12, 3d
Data Pipeline Setup :active, 2025-11-12, 3d
ML Model Skeleton & Forward Pass :2025-11-14, 4d
Evaluation Harness (Mock Metrics):2025-11-15, 3d
Status Report & Risk Assessment :milestone, 2025-11-18, 1d
section Week 3: Integration & Iteration
Backend <-> Frontend Integration :2025-11-19, 4d
Full Recommendation Pipeline Test:2025-11-20, 4d
Train Model & Monitor Loss :2025-11-19, 5d
Offline Metrics & Sanity Checks :2025-11-22, 3d
Status Report & Risk Reassessment:milestone, 2025-11-25, 1d
section Week 4: Startup Showcase & VC Fund
Finalize Demo & Deployment :2025-11-26, 3d
Prepare Slides & Visuals :2025-11-26, 3d
Rehearse Pitch & Live Demo :2025-11-29, 2d
Startup Showcase Presentation :milestone, 2025-12-02, 1d
ROOT:.
├───.venv # Environnement virtuel Python (isolant les dépendances)
├───backend # Code source de l'API REST (FastAPI)
│ ├───app # Logique applicative (modèles, schémas, config, routers)
│ └───static # Fichiers statiques (assets, CSS, JS) de l'interface utilisateur
├───data # Conteneur des données
│ ├───clean_data # Données transformées et prêtes à l'emploi (après ETL)
│ └───raw_data # Données brutes originales (fichiers sources, extractions)
├───eval # Scripts, notebooks ou fichiers pour l'évaluation des modèles ou des pipelines
├───gcp # Scripts et configurations spécifiques aux services Google Cloud
│ ├───big_query # Code lié à l'ETL et au traitement BigQuery
│ ├───embedding # Logiciel Worker pour le calcul des embeddings ML
│ ├───postgre # Scripts de création/gestion de la base Cloud SQL (PostgreSQL)
│ └───storage # Jobs liés à GCS (Loaders, Scrapers)
├───sandbox # Espace de travail temporaire, POC (Proof of Concept) ou brouillons de code
├───wair_app.egg-info # Métadonnées de l'application (générées par Python/setuptools)
└───wair_shared # Librairie de code métier partagé (Modèles BDD, fonctions ML)
└───ml # Utilitaires ou classes pour les modèles de Machine Learning
-
The raw data ingestion process is managed by a separate pipeline responsible for extracting information from external sources.
-
This raw data is then loaded into GCS and immediately on the Bronze layer of BigQuery, specifically into the
bronze_catalog.raw_data table, preserving the original format. -
Each record is enriched with an ingestion timestamp (
_ingestion_timestamp), which is critical for enabling subsequent incremental processing by downstream ETL jobs. -
This initial step ensures that all raw data is centralized, traceable, and available on Google Cloud Platform before any data transformation takes place.
-
The Cloud Run Job extracts new data from the Bronze layer and initiates the cleaning and structuring process.
-
It normalizes data types (e.g., converting price to a decimal type and ensuring scraped_at timestamps are UTC).
-
A crucial deduplication step is performed, keeping only the latest version of each product (SKU) for the
silver_catalog.product table. -
Simultaneously, multiple image URLs are pivoted into individual rows, enriched with tracking metadata (image_id, status: PENDING), and prepared for the
silver_catalog.imagetable and the next GCS upload pipeline.
-
This pipeline processes image records previously marked 'OK' in the Silver layer, retrieving them in batches for efficient memory management. The downloaded images are fed into specialized Machine Learning (ML) models for enrichment:
-
A
Fashion CLIPmodel calculates vector embeddings (numerical representations of the image content) for similarity searching. -
A
QWEN2LLM generates descriptive text and extracts relevant metadata (like style, gender, etc.) to enrich the image record.
-
-
The calculated embeddings are normalized and, along with the generated descriptions, are then inserted into the OLTP PostgreSQL database as outfit records.
This step transforms static images into valuable, searchable vectors and descriptive records, enabling the final recommendation engine.
This project uses uv to manage dependencies and the environment.
- Create a virtual environment using
uv sync - Create a
.envfile for the environment variables. - Launch the backend using
python ./main.pywhile withbackend/as current cwd