Game4all/wair-app

Repository for the recommander systems project source code

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

wair-app

Repository for the recommander systems project source code

Gantt

gantt
    title Outfit Recommendation App - 4 Week Sprint
    dateFormat  YYYY-MM-DD
    axisFormat  %d-%b

    section Week 1: Blueprint & Pitch
    Team Assignment & Roles           :done, 2025-11-04, 1d
    Define Problem & Target Users    :done, 2025-11-04, 2d
    Choose Dataset & Data Strategy   :done, 2025-11-05, 3d
    Model Innovation Selection       :done, 2025-11-07, 2d
    Technical Stack & Wireframes     :done, 2025-11-09, 2d
    Project Plan PDF (Deliverable)   :milestone, 2025-11-11, 1d

    section Week 2: First Build
    Backend Skeleton (FastAPI)       :active, 2025-11-12, 3d
    Frontend MVP (Streamlit)        :active, 2025-11-12, 3d
    Data Pipeline Setup              :active, 2025-11-12, 3d
    ML Model Skeleton & Forward Pass :2025-11-14, 4d
    Evaluation Harness (Mock Metrics):2025-11-15, 3d
    Status Report & Risk Assessment  :milestone, 2025-11-18, 1d

    section Week 3: Integration & Iteration
    Backend <-> Frontend Integration :2025-11-19, 4d
    Full Recommendation Pipeline Test:2025-11-20, 4d
    Train Model & Monitor Loss       :2025-11-19, 5d
    Offline Metrics & Sanity Checks  :2025-11-22, 3d
    Status Report & Risk Reassessment:milestone, 2025-11-25, 1d

    section Week 4: Startup Showcase & VC Fund
    Finalize Demo & Deployment       :2025-11-26, 3d
    Prepare Slides & Visuals         :2025-11-26, 3d
    Rehearse Pitch & Live Demo       :2025-11-29, 2d
    Startup Showcase Presentation    :milestone, 2025-12-02, 1d
Loading

VERSION PYTHON utilisé : 3.10.11

📁 Structure du projet - backend

ROOT:. 
├───.venv                 # Environnement virtuel Python (isolant les dépendances)
├───backend               # Code source de l'API REST (FastAPI)
│   ├───app               # Logique applicative (modèles, schémas, config, routers)
│   └───static            # Fichiers statiques (assets, CSS, JS) de l'interface utilisateur
├───data                  # Conteneur des données
│   ├───clean_data        # Données transformées et prêtes à l'emploi (après ETL)
│   └───raw_data          # Données brutes originales (fichiers sources, extractions)
├───eval                  # Scripts, notebooks ou fichiers pour l'évaluation des modèles ou des pipelines
├───gcp                   # Scripts et configurations spécifiques aux services Google Cloud
│   ├───big_query         # Code lié à l'ETL et au traitement BigQuery
│   ├───embedding         # Logiciel Worker pour le calcul des embeddings ML
│   ├───postgre           # Scripts de création/gestion de la base Cloud SQL (PostgreSQL)
│   └───storage           # Jobs liés à GCS (Loaders, Scrapers)
├───sandbox               # Espace de travail temporaire, POC (Proof of Concept) ou brouillons de code
├───wair_app.egg-info     # Métadonnées de l'application (générées par Python/setuptools)
└───wair_shared           # Librairie de code métier partagé (Modèles BDD, fonctions ML)
    └───ml                # Utilitaires ou classes pour les modèles de Machine Learning

📥 Data Ingestion (GCS -> Bronze Layer)

  • The raw data ingestion process is managed by a separate pipeline responsible for extracting information from external sources.

  • This raw data is then loaded into GCS and immediately on the Bronze layer of BigQuery, specifically into the bronze_catalog.raw_data table, preserving the original format.

  • Each record is enriched with an ingestion timestamp (_ingestion_timestamp), which is critical for enabling subsequent incremental processing by downstream ETL jobs.

  • This initial step ensures that all raw data is centralized, traceable, and available on Google Cloud Platform before any data transformation takes place.

🧼 Cleaning and Structuring (Silver Layer)

  • The Cloud Run Job extracts new data from the Bronze layer and initiates the cleaning and structuring process.

  • It normalizes data types (e.g., converting price to a decimal type and ensuring scraped_at timestamps are UTC).

  • A crucial deduplication step is performed, keeping only the latest version of each product (SKU) for the silver_catalog.product table.

  • Simultaneously, multiple image URLs are pivoted into individual rows, enriched with tracking metadata (image_id, status: PENDING), and prepared for the silver_catalog.image table and the next GCS upload pipeline.

🤖 Image Embedding and Enrichment

  • This pipeline processes image records previously marked 'OK' in the Silver layer, retrieving them in batches for efficient memory management. The downloaded images are fed into specialized Machine Learning (ML) models for enrichment:

    • A Fashion CLIP model calculates vector embeddings (numerical representations of the image content) for similarity searching.

    • A QWEN2 LLM generates descriptive text and extracts relevant metadata (like style, gender, etc.) to enrich the image record.

  • The calculated embeddings are normalized and, along with the generated descriptions, are then inserted into the OLTP PostgreSQL database as outfit records.

This step transforms static images into valuable, searchable vectors and descriptive records, enabling the final recommendation engine.

🚀 Installation et démarrage

This project uses uv to manage dependencies and the environment.

  1. Create a virtual environment using uv sync
  2. Create a .env file for the environment variables.
  3. Launch the backend using python ./main.py while with backend/ as current cwd

Contributors

Game4allLouisDuvSachafrftmaeldeprevillexXWiZzZXx

Issues