-
Data Preprocessing:
- Clean, transform, and encode the UNSW-NB15 dataset using a custom preprocessing pipeline.
- Handle numeric and categorical features appropriately.
-
Model Development:
- Implement and compare multiple machine learning models for intrusion detection, including:
- Decision Tree
- Random Forest
- XGBoost (with GPU acceleration)
- Artificial Neural Networks (ANN)
- Convolutional Neural Networks (CNN)
- Recurrent Neural Networks (LSTM)
- Implement and compare multiple machine learning models for intrusion detection, including:
-
Hyperparameter Optimization:
- Utilize Bayesian Optimization (via scikit‑optimize) to fine-tune model parameters for improved performance.
-
Model Evaluation:
- Evaluate models using key performance metrics: accuracy, precision, recall, F1-score, AUC-ROC, balanced accuracy, and confusion matrix analysis.
- Perform cross‑validation and manual inspection of predictions to check for overfitting and ensure reliable performance.
-
End-to-End Workflow:
- Ensure a robust workflow from raw data loading and preprocessing to model training, evaluation, and saving performance metrics for comparison.
Clone this repository to your local machine:
git clone <repository_url>
cd <repository_directory>It is recommended to use a virtual environment to manage dependencies:
Windows:
python -m venv environment
.\environment\Scripts\activatemacOS/Linux:
python3 -m venv environment
source environment/bin/activateInstall the necessary libraries using pip. You can either install them individually or use the provided requirements.txt file.
To install using pip:
pip install pandas numpy joblib scikit-learn scikit-optimize xgboost tensorflowAlternatively, if a requirements.txt file is provided, run:
pip install -r requirements.txtDownload the UNSW-NB15 dataset from its official source and place the training and testing CSV files in the appropriate data directory (e.g., data/row/).
Each model is implemented in its own script. For example:
- Decision Tree:
DT.py - Random Forest:
RF.py - XGBoost:
XGBoost.py - LSTM:
LSTM.py - CNN:
CNN.py - ANN:
ANN.py
To run a script, simply execute:
python <script_name>.pyFor example, to run the XGBoost model:
python XGBoost.py-
Preprocessing Pipeline:
The preprocessing functions are defined inpreprocessing_functions.pyand a pipeline is built and saved for reuse across different models. -
GPU Acceleration:
If you plan to use GPU acceleration (e.g., with XGBoost or TensorFlow), ensure that your system has an NVIDIA GPU with CUDA and cuDNN installed. Use the commandnvidia-smito verify your GPU status. -
VSCode Issues:
If VSCode doesn’t recognize certain libraries (liketensorflow.keras), ensure that the correct Python interpreter is selected, and consider adding the virtual environment’sLib\site-packagesdirectory to your VSCode settings.