This repository contains the source code for "Guess My Exercise," an iOS application that uses computer vision and machine learning to perform real-time human pose detection and classify physical exercises.
https://user-images.githubusercontent.com/1234567/123456789-abcdef.mp4
- Real-Time Pose Estimation: Detects human body joints from the device's camera feed in real-time.
- Exercise Classification: Uses a Core ML model to classify exercises like Jumping Jacks, Lunges, and Squats.
- Live Camera Feed: Works with both front and back cameras.
- Visual Feedback: Overlays a skeletal wireframe on the detected person to visualize the pose.
- Performance-Oriented: Built with Apple's high-performance frameworks (Vision, AVFoundation) and uses a reactive architecture (Combine) for a smooth user experience.
- Action Summary: Tracks and displays a summary of the exercises performed during a session.
- Swift & UIKit: Native iOS application development.
- AVFoundation: For capturing and processing video from the camera.
- Vision: For detecting human body poses (
VNDetectHumanBodyPoseRequest). - Core ML: For on-device inference using the
ExerciseClassifier.mlmodel. - Combine: For creating a reactive data processing pipeline from video capture to prediction.
The application is built on a clean, unidirectional data flow pipeline, ensuring that the logic is easy to follow and maintain.
graph TD
subgraph "Input"
A[Camera Feed]
end
subgraph "Processing Pipeline"
B(VideoCapture)
C{VideoProcessingChain}
D[Vision Pose Detection]
E[Core ML ExerciseClassifier]
end
subgraph "UI"
F[MainViewController]
G[Pose Skeleton Overlay]
H[Prediction Labels]
end
A --> B;
B -- Frame Publisher --> C;
C -- Image --> D;
D -- Detected Poses --> C;
C -- Pose Window --> E;
E -- Action Prediction --> C;
C -- Pose to Draw --> F;
C -- Action Prediction --> F;
F --> G;
F --> H;
VideoCapture: Configures theAVCaptureSessionand publishes a stream of video frames using Combine.VideoProcessingChain: Subscribes to the frame publisher and orchestrates the entire analysis pipeline:- It uses the Vision framework to detect human poses in each frame.
- It isolates the largest pose and converts the joint data into an
MLMultiArray. - It gathers these arrays into a "window" (a sequence over time).
- It feeds this window into the Core ML
ExerciseClassifierto get a prediction.
MainViewController: Receives the final output (the pose skeleton and the action prediction) and updates the UI accordingly.
The project is organized into modules by feature:
GuessMyExercise/App: General app setup, including theAppDelegateand asset catalogs.GuessMyExercise/Views: Contains the UI components, primarilyMainViewController,SummaryViewController, and the Storyboards.GuessMyExercise/Video Capture: Handles camera input via theVideoCaptureclass.GuessMyExercise/Video Processing Chain: The core processing pipeline logic inVideoProcessingChain.swift.GuessMyExercise/Pose: ThePosedata structure, which represents the detected skeleton, along with its components (Landmark,Connection).GuessMyExercise/Action Classifier: TheExerciseClassifier.mlmodeland Swift files for interacting with it.GuessMyExercise/Utility: Helper classes like thePerformanceReporter.
- Clone this repository.
- Open
GuessMyExercise.xcodeprojin Xcode. - Connect a physical iOS device (iPhone or iPad). The camera is required, so it will not run correctly on a simulator.
- Select your device as the run target.
- Build and run the application.
This project was created by Lavanya Jain.
The code in this repository is licensed under the MIT License.
Copyright (c) 2025 Lavanya Jain