model-runner is a Python-based CLI to submit hyper-parameter optimization (i.e. grid-search) jobs to the
LSF batch system.
model-runner requires a .json config as input and validates the content of the config to prevent common errors such
as non-existent files or incorrectly formatted LSF batch submissions parameters.
model-runner was tested on ETH Zurich's EULER cluster (IBM Spectrum LSF Standard 10.1.0.7) and may not work on other
clusters or lsf versions.
-
clone the repository
git clone https://github.com/kevinyamauchi/model-runner
-
Navigate to the cloned directory
cd model-runner -
We recommend using a virtual environment. If you are using anaconda, you can use the following.
conda create -n model-runner python=3.8
-
Activate your virtual environment, after it has been created. NOTE: Make sure the dependencies for your training routine (i.e.
runner) are installed.conda activate model-runner
-
Install
model-runnerin editable mode with the development dependencies.-
as a user
pip install . -
as a developer (in editable mode with development dependencies and pre-commit hooks)
pip install -e ".[dev]" pre-commit install
-
-
Activate the virtual environment you installed model-runner to. To stay consistent with the above installation
conda activate model-runner
-
Write your
.jsonconfig file following the config example. Our example configconfig_example.jsoncontains the following parameters-
job_parameters: Dictionary containing the-
(optional)
gpu_type: Type of GPU that is accepted by bsub -R "select[gpu_model0=={gpu_type}]". See this list of accepted GPUs. If not specified, the job will be run on CPU. -
logfile_dir: Path to directory to which lsf output files are saved. -
memory: Amount of memory requested per processor core. Corresponds tobsub -R "rusage[mem={memory}]". -
(optional)
ngpus: Number of GPUs ofgpu_typerequested. Corresponds tobsub -R "rusage[ngpus_excl_p={ngpus}]"and requiresgpu_typeto be specified. -
njobs_parallel: Number of jobs (i.e. hyper-parameter combinations) the model-runner will submit in parallel.
-
processor_cores: Number of processor cores requested for a single job. Corresponds tobsub -n {processor_cores}. -
run_time: The time resources on the compute note are reservered for a single job. Corresponds tobsub -W {run_time}. -
scratch: Amount of local scratch requested per processor core. Corresponds tobsub -R "rusage[scratch={scratch}]".
-
-
job_prefix: Prefix that precedes all results folders and experiment specific files. Will be appended tooutput_base_dirto create subfoldersjob_prefix{ID}inoutput_base_dircontaining all results of run {ID}. -
output_base_dir: Directory to which all config files and results folders are saved.output_base_dirmust be an input argument to the runner and should not be repeated in therunner_parameters. -
runner: Path to the file that trains your model. -
runner_parameters: Dictionary containing the parameters grid submitted to the runner. All values of the dictionary are of typelist(), includingdataandoutput_base_dir.-
data: a list of paths to your data set file or directory (whatever yourrunneraccepts as input data). Note:runnerhas to accept input data via{runner} --data {data-path}. -
{parameter_placeholder}: You can submit as many parameters as yourrunneraccepts input arguments (besidesdata,output_base_dir). You just have to make sure that{parameter_placeholder}matches an input argument ofrunnerand the values of{parameter_placeholder}are wrapped in a list. Checkout the config example for an illustrative example.- If the
type()of theparameter_placeholdervalue isbool(seeaugmentkey in example config)Truecorresponds to adding the--parameter_placeholderflag (i.e.python my_runner.py --parameter_placeholder).Falsecorresponds to omitting the flag from therunnercall (i.e.python my_runner.py). NOTE: Beware of inverted logic if parser argument usesaction=store_false.
- If the
-
-
-
Submit the hyper-parameter optimization
model_runner --params my_config.json