Benji377/datascience-template

Templates for Data Science

★ 0Forks 0PythonGitHub ↗Compare

README

Data Science Competition Snippets

A collection of reusable Python snippets for data science competitions (Kaggle, etc.).
Copy-paste these directly into notebook cells — they assume the standard Kaggle starter block with numpy, pandas, and os already imported.

Structure

snippets/
├── data_analysis/        # Explore and understand your data
│   ├── basic_eda.py            - Shape, dtypes, head/tail, describe
│   ├── missing_values.py       - Find, visualize, and handle missing data
│   ├── correlation_analysis.py - Correlation matrices and heatmaps
│   ├── distribution_plots.py   - Histograms, KDE, box plots
│   ├── target_analysis.py      - Analyze target variable distribution and balance
│   ├── outlier_detection.py    - IQR and z-score based outlier detection
│   └── feature_relationships.py- Pairplots, scatter matrices, grouped stats
│
├── models/               # Ready-to-use model templates
│   ├── logistic_regression.py  - Binary/multi-class logistic regression
│   ├── random_forest.py        - Random Forest classifier and regressor
│   ├── xgboost_model.py        - XGBoost with typical competition params
│   ├── lightgbm_model.py       - LightGBM with typical competition params
│   ├── catboost_model.py       - CatBoost (handles categoricals natively)
│   ├── knn.py                  - K-Nearest Neighbors
│   ├── svm.py                  - Support Vector Machines
│   ├── neural_net_tabular.py   - Simple neural network for tabular data
│   └── cross_validation.py     - K-Fold, Stratified K-Fold, GroupKFold
│
├── model_analysis/       # Evaluate and interpret model output
│   ├── classification_metrics.py - Accuracy, precision, recall, F1, AUC
│   ├── regression_metrics.py     - MAE, MSE, RMSE, R²
│   ├── confusion_matrix.py       - Confusion matrix visualization
│   ├── feature_importance.py     - Tree-based and permutation importance
│   ├── learning_curves.py        - Plot training vs validation curves
│   ├── prediction_analysis.py    - Residuals, error distribution
│   └── shap_analysis.py          - SHAP value explanations
│
└── utilities/            # Helpers and common operations
    ├── seed_everything.py        - Reproducibility across all libraries
    ├── feature_engineering.py    - Common feature transforms
    ├── encoding.py               - Label, one-hot, target encoding
    ├── scaling.py                - StandardScaler, MinMax, RobustScaler
    ├── memory_reduction.py       - Reduce DataFrame memory usage
    ├── timer.py                  - Time your code blocks
    └── submission.py             - Generate Kaggle submission files

Guides

The docs/ folder contains step-by-step guides for common Kaggle competition types. Start with Getting Started if this is your first competition, then pick the guide that matches your task.

Examples

The examples/ folder has full end-to-end Jupyter notebooks that use mock data to walk through a complete competition workflow — from EDA to submission file. No external data files needed.

Usage

  1. Start a new Kaggle notebook (the standard imports are already in the first cell)
  2. Browse the folder that matches what you need
  3. Copy the snippet into a new cell
  4. Adjust variable names (df, train, target, etc.) to match your data

Contributors

Benji377

Issues