ML Pipeline that allows to test different imbalance classification approaches systematically.
| Stage | Script | Purpose |
|---|---|---|
| Load CSV | src/data_loader.py | Streams the semicolon-delimited CSV in 50 k-row chunks; keeps every feat_ column; casts features to float32 and the label (Class) to int8. |
| Group split | src/split_data.py | GroupShuffleSplit (80 / 20, random_state) so each Info_group appears only in train or test. |
| Pipeline | src/pipeline.py | sampler → StandardScaler → PCA → classifier; PCA keeps 50 components (~85 % variance). Evaluates accuracy, precision, recall, F1, ROC-AUC. Saves plots. |
| Entry point | main.py | CLI / .env front-end. Lets you choose sampler, classifier, random seed and dataset path. |
Before utilizing this pipeline, ensure that you have activated a Python environment with the specified requirements:
python -m venv venv
source venv/bin/activate # Windows - .\venv\Scripts\activate
pip install -r requirements.txtTo execute this pipeline, a single entry point has been provided main.py which can be executed in two different ways:
python main.pyor:
python main.py -default_valuesIf you wish to use the -default_values flag, you must first populate the .env with the required parameters:
NAME=<pipeline name> # Used to identify the run.
FILE_PATH=<csv_file_path>
CLASSIFIER=<classifier>
IMBALANCE_HANDLER=<sampler>
RANDOM_STATE=<random_state>The pipeline will return the following information:
- Accuracy
- Precision
- Recall
- F1-Score
- ROC-AUC
- Graphs:
- confusion_matrix_<pipeline_name>.png
- precision_recall_curve_<pipeline_name>.png
- roc_curve_<pipeline_name>.png