Scot3004/etl

simple ETL example

★ 0Forks 1Jupyter NotebookGitHub ↗Compare

README

What

This repo contains script for demonstrating a simple ETL data pipeline. Starting from extracting data from the source, transforming into a desired format, and loading into a SQLite file.

Then, perform simple analysis queries on the stored data. See: analysis notebook.

Note: The used data is about the US population and unemployment rate over the past decade.


Getting started

First, make sure to have Python 3.12 installed

Then, install the project dependencies

pip install -r requirements.txt

Note: preferable, run under a virtualenv

$ python -m venv venv
$ pip install -r requirements.txt

Second, run the main pipeline file

python pipeline.py

Requirements

  • pandas
  • xlrd (for reading excel file)
  • Python >= 3.6

Data source and description

POPULATION BY METROPOLITAN AREA AND COUNTY

UNEMPLOYMENT BY COUNTY

Contributors

danielgalvis98Scot3004iamaziz

Issues