dongsukjang26/2023-summer-project

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

Word2Vec from Scratch with "Alice's Adventures in Wonderland"

This project is a Python implementation of the Word2Vec model, built entirely from scratch. The main goal is to gain a deep, practical understanding of the core mechanics behind Word2Vec without relying on high-level libraries like Gensim.

The model is trained on the classic text of "Alice's Adventures in Wonderland". It implements the Skip-gram architecture, optimized with Negative Sampling and Subsampling of frequent words.

Key Features

  • Pure Python & NumPy Implementation: The core logic for forward and backward propagation is built using only NumPy.

  • Skip-gram Model: The model learns word embeddings by predicting surrounding context words from a given target word.

  • Negative Sampling: For efficient training, the model learns to distinguish true context words from a small number of randomly selected "negative" samples, avoiding the computational bottleneck of a full softmax.

  • Subsampling: The impact of overly frequent words (e.g., 'the', 'and') is reduced by probabilistically discarding them during training. This leads to better representations for less frequent words and speeds up the process.

Project Structure

.
└── word2vec_from_scratch.ipynb  # A Jupyter Notebook containing the full implementation,
                                # data loading, training process, and evaluation.

Note: The "Alice's Adventures in Wonderland" corpus is loaded directly within the notebook, likely from a library like NLTK or a text file.

How It Works

The implementation in word2vec_from_scratch.ipynb follows these main steps:

  1. Data Loading & Preprocessing: The text of "Alice's Adventures in Wonderland" is loaded, cleaned (converted to lowercase, punctuation removed), and tokenized into a list of words. A vocabulary is then built, mapping each unique word to an integer index.

  2. Subsampling: Frequent words in the corpus are identified and down-sampled according to a probability formula to improve model performance.

  3. Generating Training Batches: The model iterates through the text. For each target word, it creates positive samples (actual words within its context window) and negative samples (random words from the vocabulary).

  4. Model Training:

    • Two weight matrices, W_in (input/target weights) and W_out (output/context weights), are initialized randomly. These matrices will become our word embeddings.

    • Using gradient descent, the model adjusts the weights to increase the similarity score for positive samples and decrease it for negative samples.

  5. Evaluation: After training, the quality of the embeddings is checked by finding the most similar words for a given test word using cosine similarity between their vectors.

Usage

To explore this project:

  1. Clone the repository to your local machine:
git clone [https://github.com/JamesJang26/2023-summer-project.git](https://github.com/JamesJang26/2023-summer-project.git)
cd 2023-summer-project
  1. Make sure you have necessary Python libraries like Numpy, collections, and Jupyter Notebook.

  2. Launch Jupyter and open word2vec_from_scratch.ipynb to run the code step-by-step and see the results.

Example Result

After training on the story, you can test the learned embeddings. For instance, finding words similar to "alice" might yield results like:

  • she

  • hatter

  • queen

  • turtle

(Note: Actual results will depend on the specific hyperparameters, training duration, and random seed.)

Contributors

dongsukjang26

Issues