nikkhn/ImageMatching

Image recommendation for unillustrated Wikipedia articles

★ 0Forks 0GitHub ↗Compare

README

ImageMatching

Image recommendation for unillustrated Wikipedia articles

Production data ETL

etl contains pyspark utilities to transform the algo raw output into a production dataset that will be consumed by a service.

spark-submit etl/transform.py <raw data> <production data>
=======
## Getting started

Connect to stat1005 through ssh (the remote machine that will host your notebooks)

ssh stat1005.eqiad.wmnet


### Installation
First, clone the repository
```shell
git clone https://github.com/clarakosi/ImageMatching.git

Setup and activate the virtual environment

cd ImageMatching
virtualenv -p python3 venv
source venv/bin/activate

Install the dependencies

export=http_proxy=http://webproxy.eqiad.wmnet:8080
export=https_proxy=http://webproxy.eqiad.wmnet:8080
python3 setup.py install

Running the script

To run the script pass in the snapshot (required), language (defaults to all wikis), and output directory (defaults to Output)

python3 algorunner.py 2020-12-28 hywiki Output

The output .ipynb and .tsv files can be found in your output directory

ls Output
hywiki_2020-12-28.ipynb  hywiki_2020-12-28_wd_image_candidates.tsv

Contributors

gmodenamirrysclarakosi

Issues