Dataset: by Mina Sameh55: https://www.kaggle.com/datasets/minasameh55/king-country-houses-aa
This project explores and models house sale prices in King County, Washington (including Seattle), covering transactions from May 2014 to May 2015.
- Clean and explore real estate data,
- Visualize relationships between housing features and prices,
- Engineer and select predictive features,
- Train and evaluate multiple regression models,
- Optimize model performance through feature engineering and hyperparameter tuning.
This project uses Python 3.9+ and the following key libraries:
pandasnumpymatplotlibseabornscikit-learnxgboost
It includes 21 features describing houses in King County and their sale prices.
Example columns:
- price: Sale price (target variable)
- sqft_living: Living area in square feet
- bedrooms, bathrooms, floors: Basic house attributes
- lat, long: Location coordinates
- yr_built, yr_renovated: Construction and renovation years
- grade, condition: House quality indicators
Main steps:
- Converted date column to datetime format
- Removed duplicate property IDs
- Verified missing and invalid values
- Identified unrealistic data (e.g., houses with 0 bathrooms or built after sale date)
- Inspected feature distributions and correlations using heatmaps and scatterplots
Key insights:
- Strong correlation between price and sqft_living, grade, and lat
- Outliers skew the price distribution — addressed via log-scaling and capping
Visualizations included:
- Scatter plots: Price vs. living area, grade
- Histograms & KDEs: Distribution of prices (normal and log-scale)
- Heatmaps: Feature correlation analysis
These helped identify redundant or highly correlated variables for removal in model training.
To simplify the model and remove redundant features:
['id', 'date', 'waterfront', 'view', 'zipcode', 'yr_renovated', 'condition', 'sqft_living15', 'sqft_lot15', 'sqft_above', 'sqft_basement', 'sale_year']
- age_at_sale = sale_year - yr_built
- Price capping at 75th percentile (Q3) to reduce outlier influence
- Feature scaling for continuous variables (sqft_living, yr_built)
- Tested feature reintroductions like waterfront and view for model improvement
Four baseline regression models were implemented and compared:
- Linear Regression
- Random Forest Regressor Ensemble
- XGBoost Regressor
- Gradient Boosting Regressor
Evaluation metrics:
- R² Score (coefficient of determination)
- Mean Squared Error (MSE)
- Mean Absolute Error (MAE)
- Added engineered feature age_at_sale
- Applied price capping (price_cap)
- Simplified model by removing low-importance variables
- Re-tested inclusion of waterfront and view
- Standardized sqft_living and yr_built for scale alignment
used RandomizedSearchCV with cross-validation to optimize key XGBoost parameters:
- n_estimators, max_depth, learning_rate
- subsample, colsample_bytree
- regularization terms (reg_alpha, reg_lambda)
Model R² Train R² Test Linear Regression 0.67 0.65 Random Forest 0.83 0.79 XGBoost 0.94 0.88 Gradient Boost 0.86 0.83
Best performing model: XGBoost with tuned hyperparameters — robust generalization and lowest test error.
- Heatmaps: Correlations among numerical features
- Price distributions: Raw and log-transformed
- Feature importance plots: From XGBoost (by gain & weight)
- Connected Dot-Plots: R^2 Test/Train over Models & Iterations
- Actual vs Predicted: Scatter plot for model fit visualization
- Size, location, and grade are the strongest predictors of house price.
- Outlier control (capping/log) significantly improves model stability.
- XGBoost consistently outperforms simpler ensemble methods with tuning.
- Feature simplification can maintain accuracy while improving interpretability & computability.
Potential extensions:
- Add geospatial features (distance to center, amenities)
- Introduce temporal trends (month or season of sale)
- Notebook Code: main.ipynb
- Readme-File: README.md
- Dataset: king_country_houses_aa.csv
- Presentation: House_Price_Prediction_SL
This work was done as a group project with Hayley & Mohamad.