ItsShadowl/fl-crawler

โ˜… 1Forks 0PythonGitHub โ†—Compare

README

Freelancer Contest Image Scraper

A Python-based web scraper for downloading contest entry images from Freelancer.com. Supports pagination, full-resolution image downloads, and can connect to your existing Chrome browser session.

Features

  • ๐Ÿš€ Multi-page scraping - Automatically navigates through all contest pages
  • ๐Ÿ–ผ๏ธ Full resolution downloads - Converts thumbnail URLs to full-resolution images
  • ๐Ÿ”ง Modular architecture - Clean, maintainable code structure
  • ๐ŸŒ Remote debugging support - Connect to existing Chrome instance (no need to close your browser!)
  • โš™๏ธ Configurable - Environment variables and command-line options
  • ๐Ÿ“ฆ Context manager support - Automatic resource cleanup
  • ๐ŸŽฏ Smart filtering - Excludes tracking pixels and banner images

Installation

Prerequisites

  • Python 3.8 or higher
  • Google Chrome browser
  • ChromeDriver (automatically managed by webdriver-manager)

Setup

  1. Clone the repository:

    git clone https://github.com/ItsShadowl/fl-crawler.git
    cd fl-crawler
  2. Create a virtual environment:

    python -m venv .venv
  3. Activate the virtual environment:

    • Windows:
      .venv\Scripts\activate
    • Linux/Mac:
      source .venv/bin/activate
  4. Install dependencies:

    pip install -r requirements.txt
  5. Configure environment (optional):

    copy .env.example .env
    # Edit .env with your preferences

Usage

Quick Start (Using Existing Chrome)

  1. Start Chrome with remote debugging:

    "C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222
  2. Log in to Freelancer.com in your Chrome browser

  3. Run the scraper:

    python main.py --contest-url https://www.freelancer.com/contest/1234567

Using New Chrome Instance

python main.py --contest-url <URL> --email [email protected] --password yourpass --new-chrome

Command Line Options

--contest-url URL        Contest URL to scrape (required)
--email EMAIL           Freelancer email/username
--password PASSWORD     Freelancer password
--download-dir DIR      Download directory (default: downloads)
--use-existing-chrome   Connect to existing Chrome (default)
--new-chrome           Start new Chrome instance
--debug-port PORT      Chrome debugging port (default: 9222)
--image-width WIDTH    Image width (default: 1920)
--no-download          Only collect URLs without downloading

Examples

Scrape and download images:

python main.py --contest-url https://www.freelancer.com/contest/1234567

Only collect URLs (no download):

python main.py --contest-url <URL> --no-download

Custom download directory:

python main.py --contest-url <URL> --download-dir my_images

Different debugging port:

python main.py --contest-url <URL> --debug-port 9223

Programmatic Usage

from src import FreelancerScraper, Config

# Create custom configuration
config = Config(
    use_existing_chrome=True,
    download_dir="my_downloads",
    image_width=2560
)

# Use as context manager
with FreelancerScraper(config) as scraper:
    # Login if needed
    # scraper.login("email", "password")
    
    # Scrape contest
    urls = scraper.scrape_contest(
        "https://www.freelancer.com/contest/1234567",
        download=True
    )
    
    print(f"Downloaded {len(urls)} images")

Project Structure

fl-crawler/
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ __init__.py       # Package initialization
โ”‚   โ”œโ”€โ”€ config.py         # Configuration management
โ”‚   โ”œโ”€โ”€ driver.py         # WebDriver setup
โ”‚   โ”œโ”€โ”€ downloader.py     # Image download logic
โ”‚   โ””โ”€โ”€ scraper.py        # Main scraper class
โ”œโ”€โ”€ main.py              # CLI entry point
โ”œโ”€โ”€ index.py             # Legacy script (deprecated)
โ”œโ”€โ”€ requirements.txt     # Python dependencies
โ”œโ”€โ”€ .env.example         # Environment variables template
โ””โ”€โ”€ README.md           # This file

Configuration

Environment Variables

Create a .env file based on .env.example:

CHROME_DEBUG_PORT=9222
USE_EXISTING_CHROME=true
DOWNLOAD_DIR=downloads
IMAGE_WIDTH=1920
PAGE_LOAD_TIMEOUT=20
LOGIN_TIMEOUT=60
ELEMENT_TIMEOUT=30

Config Class

from src import Config

config = Config(
    remote_debugging_port=9222,
    use_existing_chrome=True,
    download_dir="downloads",
    image_width=1920,
    page_load_timeout=20,
    login_timeout=60,
    element_timeout=30,
    page_transition_delay=2.0,
    scroll_delay=1.0
)

How It Works

  1. Connection: Connects to existing Chrome instance via remote debugging or starts new instance
  2. Authentication: Uses existing session or performs login
  3. Navigation: Opens contest URL and waits for page load
  4. Extraction: Finds all contest entry images on current page
  5. URL Processing: Converts thumbnail URLs to full-resolution URLs
  6. Pagination: Clicks "Next" button and repeats until last page
  7. Download: Downloads all unique images with progress tracking
  8. Cleanup: Saves URL list and closes browser (if new instance)

Troubleshooting

"Chrome already running" error

  • Close all Chrome windows before running with --new-chrome
  • Or use --use-existing-chrome (default) to connect to existing instance

Login issues

  • Use existing Chrome session for best results
  • Check that email/password are correct
  • Ensure no CAPTCHA blocking login

Images not downloading

  • Check internet connection
  • Verify contest URL is correct
  • Ensure you have write permissions in download directory

Pagination not working

  • Page structure may have changed - check console output
  • Increase page_transition_delay in config

Security Notes

โš ๏ธ Never commit credentials to version control!

  • Use environment variables for sensitive data
  • Keep .env file local (it's in .gitignore)
  • Use existing Chrome session to avoid storing passwords

License

MIT License - See LICENSE file for details

Contributing

Contributions welcome! Please:

  1. Fork the repository
  2. Create a feature branch
  3. Make your changes
  4. Add tests if applicable
  5. Submit a pull request

Support

For issues and questions:

  • Open an issue on GitHub
  • Check existing issues(if any) for solutions
  • Provide contest URL and error messages when reporting bugs

Changelog

Version 1.0.0

  • Initial modular release
  • Remote debugging support
  • Full pagination support
  • CLI interface
  • Context manager support
  • Comprehensive configuration options

Contributors

ItsShadowl

Issues