A collection of machine learning algorithms implemented from scratch using Python and NumPy, with the goal of understanding the mathematical intuition and internal working of commonly used ML algorithms without relying on high-level machine learning libraries.
The repository focuses on learning how these algorithms work internally, including model training, prediction, optimization, and core mathematical concepts.
Implementation of Linear Regression from scratch, including:
- Linear Regression using the analytical approach
- Linear Regression using Gradient Descent
- Cost function and parameter optimization
- Model prediction and evaluation
Implementation of the K-Nearest Neighbors algorithm for classification.
Key concepts:
- Distance-based classification
- Finding nearest neighbors
- Majority voting
- Choice of
k - Iris dataset example
Implementation of a Decision Tree classifier from scratch.
Key concepts:
- Entropy
- Information Gain
- Feature selection
- Best split selection
- Recursive tree construction
- Stopping criteria
- Majority-class prediction
Implementation of Random Forest using multiple Decision Trees.
Key concepts:
- Bootstrap sampling
- Multiple Decision Trees
- Ensemble learning
- Majority voting for classification
- Mean prediction for regression
- Reducing overfitting through ensemble learning
ML-Algorithms-from-Scratch/
│
├── DecisionTree/
│ ├── DecisionTree.py
│ ├── decision_tree.ipynb
│ └── train.py
│
├── RandomForest/
│ ├── DecisionTree.py
│ ├── RandomForest.py
│ ├── RandomForest.ipynb
│ └── train.py
│
├── KNN/
│ ├── KNearestNeighbour.ipynb
│ ├── Iris.csv
│ └── .gitkeep
│
├── Linear Regression/
│ ├── LinearRegression.ipynb
│ ├── LinearRegression using Gradient Descent.ipynb
│ └── Image.png
│
├── data/
│
├── .gitignore
└── LICENSE
- Python
- NumPy
- Pandas
- Matplotlib
- Jupyter Notebook
This repository is developed to build a strong understanding of machine learning algorithms by implementing them from the ground up.
The main objectives are:
- Understand the mathematical foundations of ML algorithms.
- Implement algorithms without relying on high-level ML libraries.
- Understand how training and prediction work internally.
- Study optimization and model-building techniques.
- Develop a stronger foundation for implementing more advanced machine learning systems.
- Linear Regression
- Linear Regression using Gradient Descent
- K-Nearest Neighbors
- Decision Tree
- Random Forest
- Logistic Regression
- Naive Bayes
- Support Vector Machine
- K-Means Clustering
- Principal Component Analysis
- Gradient Boosting
- XGBoost
This project is licensed under the MIT License.