hardness1020/Medical_Data_Science_Tutorial

โ˜… 10Forks 0Jupyter NotebookGitHub โ†—Compare

README

Medical Data Science Tutorial

๐Ÿ“ Table of Contents


๐Ÿง About

Aim

A tutorial for establishing an ensemble model to stratify patients with acute myeloid leukemia (AML) into three risk groups based on tabular non-clinician-initiated data.

The performance of the model is compared to the European LeukemiaNet (ELN) risk stratification system.

Dataset

The tabular data used in this tutorial is from the paper: Unified classification and risk-stratification in Acute Myeloid Leukemia.

  • NCRI cohort (for training and validation)
  • SG cohort (for testing, external cohort).

๐Ÿ Getting Started

If you don't have GPU, try using Google Colab.

Installing (local usage)

Install the packages

pip3 install -r requirements.txt

๐Ÿ”ง Content

The tutorial is divided into three parts:

1. Data Preprocessing

  • Feature selection
  • Data normalization

2. Model training

  • Model selection: random forest, xgboost
  • Hyperparameter optimizer: hyperopt
  • Ensemble: loss-based weighting

3. Model evaluation

  • Performance metrics: Accuracy, F1-score
  • Visualization: Confusion matrix
  • Survival Analysis: Kaplan-Meier estimator, C-index

๐ŸŽˆ Usage

Folder Structure

โ”œโ”€โ”€ dataset : row data
โ”‚   โ”œโ”€โ”€ NCRI.tsv : NCRI cohort (for training and validation)
โ”‚   โ””โ”€โ”€ SG.tsv   : SG cohort (for testing, external cohort)
โ”‚
โ”œโ”€โ”€ utils: utility functions
โ”‚   โ”œโ”€โ”€ hyperparameters
โ”‚   โ”‚   โ”œโ”€โ”€ hyperoptimizer.py : hyperparameter optimizer
โ”‚   โ”‚   โ””โ”€โ”€ space.py          : hyperparameters spaces for each model
|   โ”œโ”€โ”€ aml_spliter.py          : split and normalize the dataset into training, validation set 
|   โ”œโ”€โ”€ get_image_bytes.py      : plot and convert confusion matrix to bytes
|   โ”œโ”€โ”€ KM_survival_analysis.py : plot the survival curve and calculate the p-value
|   โ””โ”€โ”€ selected_features.py    : feature selected in the study
|   
โ”œโ”€โ”€ tutorial.ipynb: tutorial (not yet provided)
|
(The tutorial will produce the following files)
|
โ”œโ”€โ”€ data_preprocessed : preprocessed data
โ”‚   โ”œโ”€โ”€ NCRI.csv
โ”‚   โ””โ”€โ”€ SG.csv
|
โ””โ”€โ”€ ESB_result
    โ”œโ”€โ”€ best_trial : store the best parameters of each model
    โ”œโ”€โ”€ models     : store each model with the best parameters and weights of each model
    โ””โ”€โ”€ prediction : store the predictions of each model
        โ”œโ”€โ”€ train      : store the predictions of the training set
        โ”œโ”€โ”€ validation : store the predictions of the validation set
        โ””โ”€โ”€ external   : store the predictions of the test set

Tutorial

Follow the step by step in tutorial.ipynb (not yet provided)


โœ๏ธ Authors


๐ŸŽ‰ Reference

Contributors

hardness1020

Issues