Image recommendation for unillustrated Wikipedia articles
etl contains pyspark utilities to transform the
algo raw output into a production dataset that will be consumed by a service.
spark-submit etl/transform.py <raw data> <production data>
=======
## Getting started
Connect to stat1005 through ssh (the remote machine that will host your notebooks)ssh stat1005.eqiad.wmnet
### Installation
First, clone the repository
```shell
git clone https://github.com/clarakosi/ImageMatching.git
Setup and activate the virtual environment
cd ImageMatching
virtualenv -p python3 venv
source venv/bin/activateInstall the dependencies
export=http_proxy=http://webproxy.eqiad.wmnet:8080
export=https_proxy=http://webproxy.eqiad.wmnet:8080
python3 setup.py installTo run the script pass in the snapshot (required), language (defaults to all wikis), and output directory (defaults to Output)
python3 algorunner.py 2020-12-28 hywiki OutputThe output .ipynb and .tsv files can be found in your output directory
ls Output
hywiki_2020-12-28.ipynb hywiki_2020-12-28_wd_image_candidates.tsv