This project was developed to use computer vision to monitor cats in real time. In addition to just monitoring them, it is able to classify cat behaviors. It boasts a built out pipeline for data collection, model training, evaluation, and live deployment.
- Real-Time Cat Detection: Utilized YOLOv8n to detect cats from a given video stream.
- Activity Classification: Using a finely tuned MobileNetV3 model to classify a cat's activity to one of four categories: eating, sleeping, playing, or idle.
- Flexible Modes:
- Live Classification Mode: Runs the detector and classifier in real time to detect cats on the feed, and classify them to one of the activities.
- Data Labeling Mode: Runs a global hot key system and the detector to allow you to easily and quickly capture and label new entries into the dataset. The hot keys work even when the detector window is not in focus making it easier to label.
- Complete Pipeline
train.py: A thorough script to train the MobileNet model for activity classification using transfer learning and a two phase fine tuning approach.evaluate.py: Generates a confusion matrix to gain insight into where the model is succeeding, and failing. This helps to analyze performance, and aid in identifying areas needing improvement.review.py: An interactive tool to quickly and efficiently clean up datasets. This helped in reducing ambiguity and increasing model accuracy.
- Programming Language: Python
- Computer Vision: OpenCV
- Deep Learning Framework: PyTorch
- Object Detection Model: YOLOv8n
- Classification Model: MobileNetV3
- Utilities: 'pynput' for global keyboard listening, 'scikit-learn' for evaluation metrics, 'seaborn' and 'matplotlib' for matrix plotting.
-
Clone the repo:
git clone https://github.com/tylerlv3/catMonitor -
Create Virtual Env (Optional):
python3 -m venv venvsource venv/bin/activate -
Install Requirements:
pip install -r requirements.txt -
Download YOLO weights: The
detector.pyscript will attempt to download the YOLOv8n weights upon first run.
The project has several scripts that do different tasks to complete the ML creation process.
First, you must collect some high quality data for the respective categories.
-
To manually label images from a livestream or screen stream:
- In
main.py, setMANUAL_CLASSIFICATION = True. - Run
main.py,python main.py - a window will appear with the feed in it whether a camera or your screen. It is now listening for global key presses.
- pressing the keys (denoted in the top left of the feed screen):
e: eating,s: sleeping,p: playing,i: idle. Pressing any of these will save the last detection and label it with the key pressed. - press
q: exit when done labeling.
- In
-
To review and clean your new dataset:
- run
review.py,python review.py. - A window will open displaying the first image contained in the first directory of the dataset. You will then be taken through all images sequentially in both
dataset/trainanddataset/val. On each image you have a choice of keys to press:kwill keep the image,dwill delete the image. Once all images are decided on, the window will close.
- run
Once you have acquired some good data, you can move onto training the model.
- Run the training script:
python train.py - This will use the images in the
dataset/directory to fine tune the model and save the best performing epoch ascat_classifier.pth.
This will analyze the performance of the newly trained model.
- Run the evaluation script:
python evaluate.py - This will generate a png of a `confusion_matrix.png' and print a detailed classification report to console.
Here is an example of a generated confusion matrix:
Use the newly trained model to classify cat behavior in real time!
- In
main.py, setMANUAL_CLASSIFICATION = False. - Run the script
main.py,python main.py - The application will open in a window similar to the manual classification one, but instead it will now draw bounding boxes around detected cats in real-time with a label of the predicted activity.
- Press
qto quit.
This project is licensed under the MIT License - see the LICENSE file for details.
