A Python-based web scraper for downloading contest entry images from Freelancer.com. Supports pagination, full-resolution image downloads, and can connect to your existing Chrome browser session.
- ๐ Multi-page scraping - Automatically navigates through all contest pages
- ๐ผ๏ธ Full resolution downloads - Converts thumbnail URLs to full-resolution images
- ๐ง Modular architecture - Clean, maintainable code structure
- ๐ Remote debugging support - Connect to existing Chrome instance (no need to close your browser!)
- โ๏ธ Configurable - Environment variables and command-line options
- ๐ฆ Context manager support - Automatic resource cleanup
- ๐ฏ Smart filtering - Excludes tracking pixels and banner images
- Python 3.8 or higher
- Google Chrome browser
- ChromeDriver (automatically managed by webdriver-manager)
-
Clone the repository:
git clone https://github.com/ItsShadowl/fl-crawler.git cd fl-crawler -
Create a virtual environment:
python -m venv .venv
-
Activate the virtual environment:
- Windows:
.venv\Scripts\activate
- Linux/Mac:
source .venv/bin/activate
- Windows:
-
Install dependencies:
pip install -r requirements.txt
-
Configure environment (optional):
copy .env.example .env # Edit .env with your preferences
-
Start Chrome with remote debugging:
"C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222
-
Log in to Freelancer.com in your Chrome browser
-
Run the scraper:
python main.py --contest-url https://www.freelancer.com/contest/1234567
python main.py --contest-url <URL> --email [email protected] --password yourpass --new-chrome--contest-url URL Contest URL to scrape (required)
--email EMAIL Freelancer email/username
--password PASSWORD Freelancer password
--download-dir DIR Download directory (default: downloads)
--use-existing-chrome Connect to existing Chrome (default)
--new-chrome Start new Chrome instance
--debug-port PORT Chrome debugging port (default: 9222)
--image-width WIDTH Image width (default: 1920)
--no-download Only collect URLs without downloading
Scrape and download images:
python main.py --contest-url https://www.freelancer.com/contest/1234567Only collect URLs (no download):
python main.py --contest-url <URL> --no-downloadCustom download directory:
python main.py --contest-url <URL> --download-dir my_imagesDifferent debugging port:
python main.py --contest-url <URL> --debug-port 9223from src import FreelancerScraper, Config
# Create custom configuration
config = Config(
use_existing_chrome=True,
download_dir="my_downloads",
image_width=2560
)
# Use as context manager
with FreelancerScraper(config) as scraper:
# Login if needed
# scraper.login("email", "password")
# Scrape contest
urls = scraper.scrape_contest(
"https://www.freelancer.com/contest/1234567",
download=True
)
print(f"Downloaded {len(urls)} images")fl-crawler/
โโโ src/
โ โโโ __init__.py # Package initialization
โ โโโ config.py # Configuration management
โ โโโ driver.py # WebDriver setup
โ โโโ downloader.py # Image download logic
โ โโโ scraper.py # Main scraper class
โโโ main.py # CLI entry point
โโโ index.py # Legacy script (deprecated)
โโโ requirements.txt # Python dependencies
โโโ .env.example # Environment variables template
โโโ README.md # This file
Create a .env file based on .env.example:
CHROME_DEBUG_PORT=9222
USE_EXISTING_CHROME=true
DOWNLOAD_DIR=downloads
IMAGE_WIDTH=1920
PAGE_LOAD_TIMEOUT=20
LOGIN_TIMEOUT=60
ELEMENT_TIMEOUT=30from src import Config
config = Config(
remote_debugging_port=9222,
use_existing_chrome=True,
download_dir="downloads",
image_width=1920,
page_load_timeout=20,
login_timeout=60,
element_timeout=30,
page_transition_delay=2.0,
scroll_delay=1.0
)- Connection: Connects to existing Chrome instance via remote debugging or starts new instance
- Authentication: Uses existing session or performs login
- Navigation: Opens contest URL and waits for page load
- Extraction: Finds all contest entry images on current page
- URL Processing: Converts thumbnail URLs to full-resolution URLs
- Pagination: Clicks "Next" button and repeats until last page
- Download: Downloads all unique images with progress tracking
- Cleanup: Saves URL list and closes browser (if new instance)
- Close all Chrome windows before running with
--new-chrome - Or use
--use-existing-chrome(default) to connect to existing instance
- Use existing Chrome session for best results
- Check that email/password are correct
- Ensure no CAPTCHA blocking login
- Check internet connection
- Verify contest URL is correct
- Ensure you have write permissions in download directory
- Page structure may have changed - check console output
- Increase
page_transition_delayin config
- Use environment variables for sensitive data
- Keep
.envfile local (it's in.gitignore) - Use existing Chrome session to avoid storing passwords
MIT License - See LICENSE file for details
Contributions welcome! Please:
- Fork the repository
- Create a feature branch
- Make your changes
- Add tests if applicable
- Submit a pull request
For issues and questions:
- Open an issue on GitHub
- Check existing issues(if any) for solutions
- Provide contest URL and error messages when reporting bugs
- Initial modular release
- Remote debugging support
- Full pagination support
- CLI interface
- Context manager support
- Comprehensive configuration options