An end-to-end analytics engineering project that processes user search behavior and content interaction data, models analytics marts in DuckDB, and exposes insights through SQL, visualizations, and a Streamlit dashboard.
Key questions to answer:
- How stable user interests are over time
- Where and how interests shift between categories
- How engaged users are based on actual activity, not assumptions
This project implements a complete analytics workflow:
- PySpark ETL for scalable, incremental processing
- LLM-based keyword classification to infer search intent
- DuckDB analytics warehouse for fast, reproducible querying
- SQL analytics views and marts as the single source of truth
- Static Python visualizations for documented outputs
- Streamlit dashboard for interactive exploration
All business logic is defined in SQL marts. Dashboards and charts are read-only consumers.
Raw Logs (Search & Content)
│
▼
PySpark ETL (incremental)
│
▼
DuckDB Warehouse
├── fact_search_behavior
├── fact_content_interactions
├── analytics_* views
└── analytics marts
│
▼
Visualizations & Streamlit Dashboard
✔ Scalable ETL with PySpark
- User-level aggregations
- Incremental processing by date
- Config-driven via config.yaml
✔ LLM-based Search Classification
- Groq LLM used to classify search keywords into genres
- Robust parsing and LLM mocking in tests
✔ Analytics Warehouse & Marts
-
DuckDB used as a lightweight analytics warehouse
-
SQL marts define metrics such as:
- search stability
- interest transitions
- engagement segmentation
✔ Correct Engagement Metrics
- ActiveDays computed from real activity
- Engagement tiers derived from observed usage
✔ Multiple Output Layers
- Static charts (PNG)
- Interactive Streamlit dashboard
- SQL marts reusable by BI tools
This project produces documented, reproducible analytics outputs.
📊 See: docs/OUTPUT.md
Includes:
- final charts
- SQL queries used
- interpretation of results
A Streamlit dashboard provides interactive access to the same analytics marts.
customer-behavior-etl/
├── config.yaml
├── src/
│ ├── etl/ # Spark ETL pipelines
│ ├── llm/ # LLM classification logic
│ ├── warehouse/ # DuckDB + SQL marts
│ ├── viz/ # Static visualization scripts
│ └── dashboard/ # Streamlit app
├── docs/ # OUTPUT.md + charts
├── tests/
├── requirements.txt
├── Dockerfile
└── README.md
git clone <repo-url>
cd customer-behavior-etl
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtAll pipeline settings are stored in config.yaml:
base_path: data/log_search
output_path: output/
max_keywords: 100
batch_size: 30
write_output: falseThis project uses Groq LLM for keyword classification.
macOS / Linux:
export GROQ_API_KEY="your_groq_api_key_here"Windows:
setx GROQ_API_KEY "your_groq_api_key_here"# Run ETL
python -m src.etl.ETL_log_search
python -m src.etl.ETL_log_content
# Load data into DuckDB
python -m src.warehouse.load_parquet
# Create analytics views & marts
python -m src.warehouse.run_sql
# Generate static charts
python src/viz/plot_search_stability.py
python src/viz/plot_interest_change_summary.py
python src/viz/plot_contract_engagement.py
PYTHONPATH=. streamlit run src/dashboard/app.pypytest -q- PySpark
- DuckDB
- SQL
- Python (matplotlib, Streamlit)
- Groq LLM