This project performs sentiment analysis on product reviews using machine learning algorithms β Naive Bayes and Random Forest. It classifies textual reviews into positive or negative sentiments based on natural language processing (NLP) and supervised learning.
- Naive Bayes Classifier: A probabilistic model based on Bayesβ Theorem, efficient for text classification.
- Random Forest Classifier: An ensemble model that builds multiple decision trees and combines their output for better accuracy and robustness.
- Preprocessing of textual data (cleaning, tokenization, stopword removal, lemmatization)
- Feature extraction using TF-IDF
- Model training and evaluation
- Comparison of Naive Bayes and Random Forest performance
- Performance metrics: Accuracy, Precision, Recall, F1-score
- Source: Kaggle
- Attributes:
Review Text,Sentiment Label (Positive/Negative)
- Python
- Pandas, NumPy
- Scikit-learn
- NLTK
- Matplotlib / Seaborn (optional for visualization)
- Convert text to lowercase
- Remove punctuation and special characters
- Tokenize text into words
- Remove stopwords
- Lemmatize words
- Apply TF-IDF vectorization
- Clone the Repository
git clone https://github.com/yourusername/sentiment-analysis-ml.git cd sentiment-analysis-ml