The dataset that we generated can be found - https://drive.google.com/file/d/1IJZ5Ao82KaXwluxrpxXpX1wwEiTpqGaD/view?usp=sharing
Steps to run code:
Pre-requisites:
- Install Spark
- Install Pyspark
- Install the necessary missing libraries like findspark
- Download all required dataset files from: http://millionsongdataset.com/
Run locally
Run MapReduce-Based CF by:
- spark-submit map_reduce.py
Run ALS-Based CF by:
- spark-submit finalALSModel.py
Run Friend-Based CF by:
- python friendBasedCF.py
You can use additional options while submitting the spark job such as --master local[8] where 8 can be replaced with the number of cores you want to use.
Run on Cloud
- Sign up on AWS services and opt for EMR and S3 services.
- Create a S3 bucket to store input/output/log data files.
- Go to EMR and create a new cluster (You can choose any number of nodes in your cluster). Additional configs related to driver/executor memory, session timeout etc. can be added as a JSON config while creating the cluster. (This can be set manually even after the cluster is started using %configure signature)
- Create a new Jupyter Notebook and link it to the newly created cluster.
- Wait for the cluster to complete set up and then start/open the Jupyter Notebook.
- Select a pySpark Kernel once it starts.
- Finally, run the Jupyter Notebook.
Extracting each audio attribute from HDF5 file
Converting the files in HDF5 format to CSV format and writing to SongCSV.csv file
Joining the data of metadata.csv and triplets.txt files over song_id to obtain consolidated dataset.
Cleaning the data and converting the tab separated dat file into csv files.
Merging the triplets dataset which consists of the userID, songID and listen count, with metadata such as song name, artist name, etc.
Uses cross validator and grid search to run the ALS model multiple times with different parameters to find the optimal hyperparameters.
"Alternating least squares (ALS)" is a distributed matrix factorization method that allows for faster and more efficient computations. The algorithm predicts number of listens using Matrix Factorization. It is an iterative algorithm that alternates back and forth between user and song vectors for solving and providing the recommendations.
Uses the best hyperparameters obtained for ALS to run the model on the entire dataset and make predictions for users.
Friend-Based CF is based on the assumption that an individuals taste/liking is strongly influence by the people around him. It is more likely that an individuals taste in music is more similar to his friend rather than a stranger in a different country. Hence, if we can define these relations and form smaller cluster then it is possible to use CF to compare an individual only to his friends and connections in order to determine similarity. Hence, a friend based collaborative filtering would be more efficient, accurate and scalable.
Creates a cluster of 2nd degree friends and calculates the cosine similarity between users of that cluster and makes recommendation. It calculates the cosine similarity of all users as well so that the results can be compared.
This approach uses item-item collaborative filtering to provide the recommendations for a user. We started with the data set consisting of user, song and rating information. We calculated all the pairs of songs listened by the users and the corresponding ratings of songs. This is done for all the users and all the songs they have listened to. Once we have this list consisting of song pairs and rating pairs, for each song pair we form a vector of ratings pairs collected by a number of users. Next the cosine similarity algorithm is applied on this vector to find the similarity score of the song pair. While providing the recommendation for a user, we consider the user’s top songs and recommend other songs which are similar to his listening history.
Create a normalized version of the listen count to check how it affects the performance of the model.
Converts the triplets file in text format into a csv format for the model to access.
Please refer to the report for additional information and experimental results.