Impact of distribution of contact by ethnicity on infectious disease dynamics in respiratory pathogens.
The code presented in this repository can be used to reproduce the analysis (using simulated contact data only) and figures presented in Robert A et al, Ethnic inequalities in respiratory virus epidemics in England: a mathematical modelling study.
- In England, is the contact distribution different between ethnic groups after adjusting for demographic variables?
- In England, how do differences in demographic distribution, mixing, and contact distribution between ethnic groups translate into unequal outbreak dynamics?
- Bayesian negative binomial regression on Reconnect social contact data
- Outcome: distribution of the number of contacts per participant (mean and dispersion).
- Variable of interest: compound variable on ethnic group and urban/rural status. Seven levels: Black ethnicity – urban, Asian ethnicity – urban, Mixed ethnicity – urban, Other ethnicity - urban, White ethnicity – urban, White ethnicity – rural, Non-white ethnicity – rural.
- Control variables: age group, income, gender, weekday, household size, employment status.
- Simulation of distribution of contacts in the population
- Generate synthetic populations (distribution of income, employment status, gender, household size per ethnic and age group following the England population).
- Simulate the number of contacts for each individual in the synthetic population using the parameters from the regression analysis.
- For each age and ethnic group, fit a mixture of three Poisson distributions to the simulated distribution of contacts to classify individuals in low, medium, and high contact groups.
- Generate stochastic simulations using a compartmental model stratified by age, ethnic, and contact group
- Use the mixing matrices from Reconnect and the low, medium, and high contact groups to compute a mixing matrix stratified by age, ethnic, and contact group.
- Simulate epidemics using a range of scenarios and conditions.
- Compare the trajectories generated by the transmission model for each ethnicity.
Clone/download this project onto your machine.
The following R packages are required to run the code:
* rio
* tidyverse
* broom.mixed
* brms
* flexmix
* grid
* ggridges
* patchwork
* odin2
* dust2
* ggpmisc
and can be installed in R by running:
install.packages("rio", "tidyverse", "broom.mixed", "brms", "flexmix", "grid", "ggridges", "patchwork", "ggpmisc")
install.packages(c("odin2", "dust2"), repos = c("https://mrc-ide.r-universe.dev", "https://cloud.r-project.org"))The code contains scripts to generate age stratified and ethnicity stratified mixing matrices from Reconnect, implement a Bayesian negative binomial regression analysis, and run stochastic simulations using a compartmental model.
First, script_contact_matrix.R can be run to compute the age-stratified and ethnicity-stratified per capita contact matrix, using the Reconnect CSV files on Zenodo, and the functions from the Reconnect Github repository. This script will generate two files in /results: dt_contact_age.RDS (number of contacts per capita between age groups) and dt_contact_eth.RDS (number of contacts per capita between ethnicities):
source("R/script_contact_matrix.R")The runtime of this script is several hours on a standard laptop with a 3.0 GHz processor and 32 GB RAM.
Then, script_regression.R runs the two Bayesian negative binomial regression models, one using only ethnicity + urban/rural status as the covariate, and one controlling for age, household size, household income, employment status, sex, and weekday of collection. Once run, the script generates a file in /results: results/regression_output_anoun.rds, which is a list of two brms files containing the parameter estimates of each model.
source("R/script_regression.R")The individual participant / contact data cannot be shared publicly, so the regression models are run on a simulated dataset. The runtime of this script does not exceed 20 minutes on a standard laptop with a 3.0 GHz processor and 32 GB RAM.
Then, script_synthetic_population.R simulates the number of contacts per individual in different scenarios, using the demographic characteristics of the population in England according to the 2021 census and the parameter estimates from the Bayesian regression model. Three scenarios are run: 1) all individuals have the same demographic characteristics (age, household size, household income, employment status, sex); only ethnicity changes, 2) the distribution of demographic characteristics for each ethnicity follows the overall distribution in the population, 3) the distribution of demographic characteristics follows the distribution in the population by ethnicity (sex, age by ethnicity, household size by age and ethnicity, household income by ethnicity, employment status by age, ethnicity, and sex). The script generates an RDS file (results/contact_distribution_synthetic_anoun.RDS), containing the number of contacts by individual in each synthetic population.
source("R/script_synthetic_population.R")The runtime of this script is under 30minutes on a standard laptop with a 3.0 GHz processor and 32 GB RAM.
Then, script_run_all_models.R generates stochastic simulation sets for each scenario based on the parameter estimates from the negative binomial regression analysis. Stochastic simulations are generated for each location and scenario presented in the paper. This script generates two RDS files: results/outputmodel_byr0_region.RDS () and results/prop_and_r0_clust_region.RDS ():
source("R/script_run_all_models.R")By default, the runtime of this script is over 30 hours on a standard laptop with a 3.0 GHz processor and 32 GB RAM. Using type <- "short" in line 3 of the script shortens the run time, but reduces the number of runs, simulations per run, and the size of the synthetic population.
Finally, script_generate_all_figures.R generates all Figures presented in the paper and supplementary material. The figures are saved in the figures/ folder.
source("R/script_generate_all_figures.R")The runtime of this script does not exceed 10 minutes on a standard laptop with a 3.0 GHz processor and 32 GB RAM.
results/contact_distribution_synthetic.RDS: Number of contacts by individual in different synthetic populations simulated from the outputs of the regression analysis.results/contact_distribution_synthetic_anoun.RDS: Number of contacts by individual in different synthetic populations simulated from the outputs of the regression analysis using anonymised contact data.results/dt_contact_age.RDS: Number of contacts per capita between age groups, generated withR/script_contact_matrix.Rusing the Reconnect outputs.results/dt_contact_eth.RDS: Number of contacts per capita between ethnicities, generated withR/script_contact_matrix.Rusing the Reconnect outputs.results/outputmodel_byr0_region.RDS: Data frame generated inR/script_run_all_models.R, contains the simulations used to compare incidence across UK cities and to show the impact of demographic characteristics on pathogen transmission (proportion of the population infected over the course of the outbreak for each ethnicity).results/outputmodel_byr0_region_4.RDS: Data frame generated inR/script_run_all_models.Rusing 4 contact groups (instead of 3 in the reference simulation set), contains the simulations used to compare incidence across UK cities and to show the impact of demographic characteristics on pathogen transmission (proportion of the population infected over the course of the outbreak for each ethnicity).results/prop_and_r0_clust_region.RDS: Data frame generated inR/script_run_all_models.R, contains the simulations used to compare R0 and incidence when specifying beta instead of R0 (proportion of the population infected over the course of the outbreak for each ethnicity).results/regression_output_anoun.rds: List of brms files, containing the parameter estimates of each regression model (generated withscript_regression.R).
The analysis uses various datasets. The Data folder contains the following files:
Data/age_ethnicity_economic.csv: Distribution of economic activity status by age, sex, and ethnic group at a national level. Source: 2021 Census data;age_ethnicity_economic.txtshows how the CSV file was created on the census website.Data/age_ethnicity_economic_london.csv: Distribution of economic activity status by age and ethnic group in London. Source: 2021 Census data;age_ethnicity_economic_london.txtshows how the CSV file was created on the census website. Due to data protection requirements, only the national-level data could be stratified by gender. The gender distribution of economic activity status at a local level is computed using the gender distribution at a national level (see functionclean_employinR/function_clean_pop_data.R, lines 176-190 of the script).Data/age_ethnicity_economic_la.csv: Distribution of economic activity status by age and ethnic group in Birmingham, Leicester, Liverpool, Manchester, York. Source: 2021 Census data;age_ethnicity_economic_la.txtshows how the CSV file was created on the census website. Due to data protection requirements, only the national-level data could be stratified by gender. The gender distribution of economic activity status at a local level is computed using the gender distribution at a national level (see functionclean_employinR/function_clean_pop_data.R, lines 206-221 of the script).Data/age_ethnicity_householdsize.csv: Distribution of household size by age and ethnic group in England. Source: 2021 Census data;age_ethnicity_householdsize.txtshows how the CSV file was created on the census website.Data/age_ethnicity_householdsize_london.csv: Distribution of household size by age and ethnic group in London. Source: 2021 Census data;age_ethnicity_householdsize_london.txtshows how the CSV file was created on the census website.Data/age_ethnicity_householdsize_la.csv: Distribution of household size by age and ethnic group in Birmingham, Leicester, Liverpool, Manchester, York. Source: 2021 Census data;age_ethnicity_householdsize_la.txtshows how the CSV file was created on the census website.Data/income_by_ethnicity.csv: Distribution of weekly household income by ethnicity in the UK. Weekly income is multiplied by 50, and matched to the categories of income used in the Reconnect study. Source: Family Resources Survey 2020/21, the data can be downloaded hereData/participants_anonymous.RDS: Simulated individual-level contact data mimicking the original Reconnect data.
Directory structure:
├── Data
│ ├── age_ethnicity_economic.csv
│ ├── age_ethnicity_economic.txt
│ ├── age_ethnicity_economic_la.csv
│ ├── age_ethnicity_economic_la.txt
│ ├── age_ethnicity_economic_london.csv
│ ├── age_ethnicity_economic_london.txt
│ ├── age_ethnicity_householdsize.csv
│ ├── age_ethnicity_householdsize.txt
│ ├── age_ethnicity_householdsize_la.csv
│ ├── age_ethnicity_householdsize_la.txt
│ ├── age_ethnicity_householdsize_london.csv
│ ├── age_ethnicity_householdsize_london.txt
│ ├── income_by_ethnicity.csv
│ └── participants_anonymous.RDS
├── R
│ ├── function_clean_pop_data.R
│ ├── function_data_cleaning.R
│ ├── function_figures.R
│ ├── function_figures_main.R
│ ├── function_figures_supplement.R
│ ├── function_run_seir_model.R
│ ├── function_seir_model.R
│ ├── function_simulated_pop.R
│ ├── library_and_scripts.R
│ ├── script_contact_matrix.R
│ ├── script_generate_all_figures.R
│ ├── script_regression.R
│ ├── script_run_all_models.R
│ └── script_synthetic_population.R
├── results
│ ├── contact_distribution_synthetic.RDS
│ ├── contact_distribution_synthetic_anoun.RDS
│ ├── dt_contact_age.RDS
│ ├── dt_contact_eth.RDS
│ ├── outputmodel_byr0_region.RDS
│ ├── outputmodel_byr0_region_4.RDS
│ ├── prop_and_r0_clust_region.RDS
│ └── regression_output_anoun.rds
└── README.md
The R folder contains various scripts and function files:
library_and_scripts.R: Lists all libraries and local files that must be imported to run the scripts.function_clean_pop_data.R: Functions to import and clean datasets on the distribution of household size, household income, and economic status in the population.function_data_cleaning.R: Function to import the contact data (only the version on simulated anonymised data is shared in this repository).function_figures.R: Functions to visualise regression model outputs, the distribution of simulated contacts in synthetic populations, and outbreak trajectories.function_figures_main.R: Functions to generate the figures in the main section of the paper.function_figures_supplement.R: Functions to generate the figures in the Supplementary section of the paper.function_simulated_pop.R: Functions to create the synthetic populations and simulate the number of contacts using the regression outputs.function_seir_model.R: SEIR compartmental model created using odin2.function_run_seir_model.R: Functions to set up and run a set of simulations using the SEIR transmission model defined infunction_seir_model.R.script_contact_matrix.R: Compute the age-stratified and ethnicity-stratified per capita contact matrix, using the Reconnect CSV files on Zenodo, and the functions from the Reconnect Github repository. This script will generate two files in/results:dt_contact_age.RDS(number of contacts per capita between age groups) anddt_contact_eth.RDS(number of contacts per capita between ethnicities).script_generate_all_figures.R: Generates all Figures presented in the paper and supplementary material and saves them in thefigures/folder.script_regression.R: Runs the two Bayesian negative binomial regression models, one using only ethnicity + urban/rural status as covariate, one controlling for age, household size, household income, employment status, sex, and weekday of collection.script_run_all_models.R: Generates stochastic simulation sets for each scenario based on the parameter estimates from the negative binomial regression analysis.script_synthetic_population.R: Simulates the number of contacts per individual in different scenarios, using the demographic characteristics of the population in England according to the 2021 census and the parameter estimates from the Bayesian regression model.