mara-pn/project-supervised-ML

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

House Price Prediction in King County

Dataset: by Mina Sameh55: https://www.kaggle.com/datasets/minasameh55/king-country-houses-aa

Project Overview

This project explores and models house sale prices in King County, Washington (including Seattle), covering transactions from May 2014 to May 2015.

  • Clean and explore real estate data,
  • Visualize relationships between housing features and prices,
  • Engineer and select predictive features,
  • Train and evaluate multiple regression models,
  • Optimize model performance through feature engineering and hyperparameter tuning.

1. Setup

1.1. Requirements

This project uses Python 3.9+ and the following key libraries:

  • pandas
  • numpy
  • matplotlib
  • seaborn
  • scikit-learn
  • xgboost

1.2. Data

It includes 21 features describing houses in King County and their sale prices.

Example columns:

  • price: Sale price (target variable)
  • sqft_living: Living area in square feet
  • bedrooms, bathrooms, floors: Basic house attributes
  • lat, long: Location coordinates
  • yr_built, yr_renovated: Construction and renovation years
  • grade, condition: House quality indicators

2. Data Exploration & Cleaning

Main steps:

  • Converted date column to datetime format
  • Removed duplicate property IDs
  • Verified missing and invalid values
  • Identified unrealistic data (e.g., houses with 0 bathrooms or built after sale date)
  • Inspected feature distributions and correlations using heatmaps and scatterplots

Key insights:

  • Strong correlation between price and sqft_living, grade, and lat
  • Outliers skew the price distribution — addressed via log-scaling and capping

3. Exploratory Data Analysis (EDA)

Visualizations included:

  • Scatter plots: Price vs. living area, grade
  • Histograms & KDEs: Distribution of prices (normal and log-scale)
  • Heatmaps: Feature correlation analysis

These helped identify redundant or highly correlated variables for removal in model training.

4. Feature Selection & Engineering

4.1 Dropped columns

To simplify the model and remove redundant features:

['id', 'date', 'waterfront', 'view', 'zipcode', 'yr_renovated', 'condition', 'sqft_living15', 'sqft_lot15', 'sqft_above', 'sqft_basement', 'sale_year']

4.2 Engineered features

  • age_at_sale = sale_year - yr_built
  • Price capping at 75th percentile (Q3) to reduce outlier influence
  • Feature scaling for continuous variables (sqft_living, yr_built)
  • Tested feature reintroductions like waterfront and view for model improvement

5. Modeling

Four baseline regression models were implemented and compared:

  • Linear Regression
  • Random Forest Regressor Ensemble
  • XGBoost Regressor
  • Gradient Boosting Regressor

Evaluation metrics:

  • R² Score (coefficient of determination)
  • Mean Squared Error (MSE)
  • Mean Absolute Error (MAE)

6. Model Optimization

6.1 Feature Engineering Iterations

  • Added engineered feature age_at_sale
  • Applied price capping (price_cap)
  • Simplified model by removing low-importance variables
  • Re-tested inclusion of waterfront and view
  • Standardized sqft_living and yr_built for scale alignment

6.2 Hyperparameter Tuning

used RandomizedSearchCV with cross-validation to optimize key XGBoost parameters:

  • n_estimators, max_depth, learning_rate
  • subsample, colsample_bytree
  • regularization terms (reg_alpha, reg_lambda)

7. Results Summary

Model R² Train R² Test Linear Regression 0.67 0.65 Random Forest 0.83 0.79 XGBoost 0.94 0.88 Gradient Boost 0.86 0.83

Best performing model: XGBoost with tuned hyperparameters — robust generalization and lowest test error.

8. Visualizations

  • Heatmaps: Correlations among numerical features
  • Price distributions: Raw and log-transformed
  • Feature importance plots: From XGBoost (by gain & weight)
  • Connected Dot-Plots: R^2 Test/Train over Models & Iterations
  • Actual vs Predicted: Scatter plot for model fit visualization

9. Key Takeaways

  • Size, location, and grade are the strongest predictors of house price.
  • Outlier control (capping/log) significantly improves model stability.
  • XGBoost consistently outperforms simpler ensemble methods with tuning.
  • Feature simplification can maintain accuracy while improving interpretability & computability.

10. Next Steps

Potential extensions:

  • Add geospatial features (distance to center, amenities)
  • Introduce temporal trends (month or season of sale)

11. Repository

  • Notebook Code: main.ipynb
  • Readme-File: README.md
  • Dataset: king_country_houses_aa.csv
  • Presentation: House_Price_Prediction_SL

12. Collaboration

This work was done as a group project with Hayley & Mohamad.

Contributors

mara-pn

Issues