- Download and extract the LexisNexisCrawler, either via https://gitlab.com/TUHH-TIE/LexisNexisCrawler/repository/archive.zip?ref=master or clone the Git repository.
- Install firefox
- Install the requirements:
python3 -m pip install -r requirements.txt - Install the crawler:
python3 setup.py install(on linux, prefixsudo) - Add your user data:
python3 configure_user_data.py, the script will guide you through it. - Follow http://stackoverflow.com/a/40208762/3212182
- Export your queries to a .csv file. Your query has to be in a column with the
title
name. Other recognized column names arefrom date,to date(in the formatDD-MM-YYYY) andlanguages(one ofus(default)all,germanorenglish, onlyusis well tested though, others might yield errors and need adaption). - Simply invoke
crawl_nexis.py <your_csv_file.csv> <output_dir> - Your data will be fetched and written into the output directory you specified as JSON files.
An example CSV file is given with example.csv. The test.csv contains some
larger query set useful for debugging.
Some queries are resulting in a large number of datasets. Those are by default ignored and will not be reattempted. Such a query will result in a JSON file like this:
{
"error_code": 1,
"error": "Too many results (>3000)"
}So if you want to further want to work with the data you can just look into the JSON error code to determine if the data was downloaded correctly.
If you want the crawler to download even big datasets, pass the -b argument.