Shalsh23/VocabBuilder

Multi-source vocabulary learning application - aggregates words from various sources for comprehensive language learning

★ 0Forks 0PythonGitHub ↗Compare

README

https://www.wordotter.com/

VocabBuilder - Multi-Source Vocabulary Application

A comprehensive vocabulary learning application that aggregates interesting words, definitions, and usage examples from multiple sources to build a rich language learning database.

Overview

VocabBuilder is designed to be an extensible vocabulary aggregation platform. Currently, it supports scraping from the "A Word A Day" (AWAD) website at wordsmith.org, with plans to integrate multiple vocabulary sources including dictionary APIs, literature databases, and other educational resources.

Vision

Some people collect stamps or pressed flowers. We collect words. Not in haste, but in the way you might collect seashells on a quiet morning—turning each one over, marveling at its shape, feeling the weight of its history.

WordOtter is for those who find joy in this kind of lingering. It's for people who read a sentence twice, not because they didn't understand it, but because it felt too good to leave behind. For those who smile when they meet sonder or halcyon in the wild, and tuck it away for later use.

We believe words are not mere instruments of communication—they are the tools with which we think, dream, and build the architecture of our inner worlds. Learning them should feel less like homework and more like stepping into a well-lit room you didn't know was in your own home.

Here you'll find:

  • A Curated Lexicon of words worth keeping
  • Precise Definitions paired with examples drawn from literature and real life
  • Interactive Study Tools that make practice a pleasure
  • Free & Open Access for anyone with curiosity and a few spare minutes

For the word-lover, this is a homecoming. For everyone else, perhaps, the beginning of an unexpected romance.

Current Features

Wordsmith.org Integration

  • Scrapes word archives from wordsmith.org/awad
  • Extracts detailed word information including:
    • Word definitions
    • Parts of speech
    • Real-world usage examples from literature and media
  • Saves data in CSV format for easy access and analysis
  • Implements polite scraping with delays to respect server resources

Web Interface (NEW!)

  • Interactive Dashboard: Browse and study vocabulary through a modern web UI
  • Word List View: Paginated display with search and sort functionality
  • Word Detail Pages: Complete information for each word with navigation
  • Advanced Search: Real-time search across words and definitions
  • Study Mode: Interactive flashcards with keyboard shortcuts
  • Responsive Design: Works on desktop, tablet, and mobile devices

Planned Features

Additional Data Sources

  • Dictionary APIs: Integration with Merriam-Webster, Oxford, Cambridge APIs
  • Literature Sources: Project Gutenberg word frequency analysis
  • Academic Sources: SAT/GRE vocabulary lists
  • Etymology Sources: Word origin and history tracking
  • Language Learning Platforms: Duolingo, Memrise vocabulary sets

Application Features

  • Unified Database: Centralized SQLite/PostgreSQL database for all vocabulary
  • Web Interface: Flask/Django web app for browsing and learning
  • API Service: RESTful API for vocabulary queries
  • Spaced Repetition: Learning algorithm for vocabulary retention
  • Daily Digest: Email/notification service for word of the day
  • Mobile App: React Native app for on-the-go learning
  • Export Options: Anki deck generation, PDF flashcards

Project Structure

Current Structure

VocabBuilder/
├── src/                     # Source code
│   ├── scrape_words.py      # Scrapes word URLs from archives
│   ├── extract_meanings.py  # Extracts detailed word information
│   └── check_status.py      # Check processing status
├── web/                     # Web interface (Flask)
│   ├── app.py              # Main Flask application
│   ├── templates/          # HTML templates
│   ├── static/             # CSS and JavaScript
│   └── README.md           # Web interface documentation
├── resources/               # Data and output files
│   ├── wordsmith_words.csv
│   ├── wordsmith_complete.csv
│   └── wordsmith_extraction.log
├── docs/                    # Documentation
│   ├── architecture.md
│   ├── features.md
│   ├── api-reference.md
│   └── roadmap.md
├── requirements.txt
├── README.md
└── .gitignore

Planned Structure

VocabBuilder/
├── scrapers/                # Data source integrations
│   ├── wordsmith/
│   ├── dictionary_apis/
│   ├── literature/
│   └── academic/
├── database/               # Database models and migrations
├── api/                    # REST API service
├── web/                    # Web interface
├── mobile/                 # Mobile application
├── utils/                  # Shared utilities
├── tests/                  # Test suite
└── data/                   # Data storage

Installation

  1. Clone the repository:
git clone [repository-url]
cd WordADay
  1. Create a virtual environment:
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate
  1. Install required packages:
pip install -r requirements.txt

Or install individually:

pip install requests beautifulsoup4 Flask

Usage

Web Interface Usage

  1. Navigate to the web directory and start the Flask server:
cd web
python app.py
  1. Open your browser and visit:
http://localhost:8080
  1. Explore the features:
    • Home: Dashboard with word count and Word of the Day
    • Words: Browse all words with pagination and search
    • Search: Advanced search with real-time results
    • Study: Flashcard mode with keyboard shortcuts (Space to flip, arrows to navigate)
    • About: Information about the project

Command-Line Usage

  1. First, scrape the word URLs from the archives:
cd src
python scrape_words.py

The scraper will automatically:

  • Check for existing words in the database
  • Report how many new words were found
  • Merge new words with existing data
  1. Then extract detailed information for each word:
python extract_meanings.py

The extractor will automatically:

  • Skip words that have already been processed
  • Resume from where it left off if interrupted
  • Save progress after each word to prevent data loss
  1. Check processing status:
python check_status.py

Resume After Interruption

Both scripts support resuming after interruption:

  • Scraper: Automatically merges new words with existing data
  • Extractor: Skips already processed words and appends new ones
  • Use Ctrl+C to safely interrupt processing at any time

Output

The project generates:

  • wordsmith_words.csv: Contains all scraped words and their URLs
  • wordsmith_complete.csv: Complete dataset with words, meanings, and usage examples
  • wordsmith_extraction.log: Detailed logs of the extraction process

Example Data

The dataset includes interesting vocabulary such as:

  • "central casting" - stereotypical
  • "bunny boiler" - a dangerously obsessive person
  • "elsewhen" - at another time
  • "towardly" - compliant or pleasant

Requirements

  • Python 3.x
  • requests (for web scraping)
  • beautifulsoup4 (for HTML parsing)
  • Flask (for web interface)

License

This project is for educational purposes. Please respect the original content source at wordsmith.org.

Author

Shalki Shrivastava

Contributors

Shalsh23

Issues