A small Python scraper that collects real-estate listings from sa.aqar.fm and exports them as a CSV file for analysis.
The scraper:
- Iterates over paginated listing pages on sa.aqar.fm
- Uses cached HTTP responses (via joblib) to avoid re-downloading the same pages
- Parses listing data primarily from embedded JSON (
__NEXT_DATA__) with a BeautifulSoup fallback - Saves all results into
data/raw/aqar_fm_listings.csvanddata/raw/aqar_fm_listings.json
⚠️ Disclaimer
This project is for personal/educational use only. When using it, you are responsible for complying with aqar.fm’s Terms of Service, robots.txt, and all applicable laws. Do not use it for abusive or high‑volume scraping.
main.py– Entry point and core scraper logicpyproject.toml– Project metadata and dependenciesdata/raw/aqar_fm_listings.csv– Output CSV (created by the scraper)data/raw/aqar_fm_listings.json– Output JSON (created by the scraper)data/processed/aqar_fm_listings_cleaned.csv– Cleaned CSV (created by the clean script)data/output/aqar_fm_listings_auction_cleaned.csv– auction CSV (all auction listings)data/output/aqar_fm_listings_rental_cleaned.csv– rental CSV (all rental listings)data/output/aqar_fm_listings_sale_cleaned.csv– sale CSV (all sale listings)data/cache/– HTTP response cache managed by joblibchecks.ipynb– Example notebook for inspecting the data (optional)
- Python 3.13+
- A working internet connection
- Basic understanding of environment variables /
.envfiles
Python dependencies (also defined in pyproject.toml):
beautifulsoup4httpxjoblibpandaspython-dotenv
-
Install uv if you don’t already have it.
-
From the project root:
uv sync
This will create a virtual environment and install all dependencies.
The site uses Cloudflare and some anti‑bot mechanisms. To make your requests look like a real browser session, you must supply a few cookies via environment variables.
Create a .env file in the project root with:
REQ_DEVICE_TOKEN=your_req_device_token_here
CF_CLEARANCE=your_cf_clearance_here
CF_BM=your_cf_bm_hereHow to obtain these:
- Open your browser’s DevTools (Network tab).
- Visit aqar.fm listings pages.
- Inspect a normal page request and copy the values of:
req-device-tokencf_clearance__cf_bm
- Paste them into
.envas shown above.
If these are missing or invalid, the script may be blocked or return “you have been blocked” in the HTML.
From the project root:
uv run main.pyThe script will:
-
Generate a list of listing URLs starting at:
rooturl = "https://sa.aqar.fm/%D8%B9%D9%82%D8%A7%D8%B1%D8%A7%D8%AA/" all_urls = [rooturl + f"{i}" for i in range(1, 9999)]
-
Fetch pages concurrently (up to 10 threads).
-
Automatically stop when it encounters a page containing the Arabic text “لا توجد نتائج” (“no results”), or if it detects that you are blocked.
-
Parse each page and extract fields such as:
title,url,price,descriptioncity,district,address,coordinates(lat,lng)sale_type(sale,rental, orauction)area_sqm,num_bedrooms,num_bathrooms,num_living_roomsfloor_level,street_width,age- Attributes:
furnished,ac,kitchen,lift,car_entrance, etc. - Media:
images,videos - Metadata:
create_time,published_at,user_info
-
Save all listings into:
data/raw/aqar_fm_listings.csv data/raw/aqar_fm_listings.json
The scraper also uses a joblib Memory cache under ./data/cache so repeated runs don’t refetch unchanged pages.
To process the raw scraped data, run:
uv run clean_data.pyThis script performs several cleaning and normalization steps:
- Deduplication: Removes duplicate listings based on ID or URL.
- Data Type Conversion: Converts prices and numeric fields (area, bedrooms, etc.) to proper number formats.
- Text Normalization:
- Normalizes Arabic text (unifying aleph forms, etc.).
- Removes diacritics (Tashkeel).
- Removes emojis and extra whitespace.
- Boolean Standardization: Converts various yes/no/1/0 formats to standard booleans.
- Dataset Splitting: Separates the data into three categories based on
sale_type:- Sale: Listings for sale.
- Rental: Listings for rent.
- Auction: Listings for auction.
Outputs:
The script generates the following files in data/processed/ and data/output/:
data/processed/aqar_fm_listings_cleaned.csv(Full cleaned dataset)data/output/aqar_fm_listings_sale_cleaned.csvdata/output/aqar_fm_listings_rental_cleaned.csvdata/output/aqar_fm_listings_auction_cleaned.csv
(JSON versions are also generated for each)
You can tweak behavior directly in main.py:
-
Starting URL / category
Changerooturlto scrape a different path or category on aqar.fm. -
Concurrency
Adjustmax_workersin theThreadPoolExecutorto control the number of concurrent requests. -
Page limit / early stop
- The script will stop automatically when it sees “لا توجد نتائج”.
- It also uses a global
STOP_PAGEto remember the first page without results. - To impose a hard limit, you can:
- Reduce the
range(1, 9999)to a smaller number of pages, or - Manually set
STOP_PAGEinmain.pyto something finite.
- Reduce the
-
Parsed fields
The CSS selectors and icon-to-field mapping live inparse_category_page().
You can extend or modify these to extract additional fields.
- If the site changes its HTML structure or CSS classes, parsing may break; in that case, update the selectors in
parse_category_page(). - If your cookies expire or change, you’ll need to refresh the
.envvalues. - High-frequency scraping might trigger additional anti-bot measures. Consider:
- Lowering concurrency
- Adding small random delays
- Running less frequently
This project is dual-licensed: