Mitigating health inequities with machine learning: a nationwide cohort study developing and evaluating ethnicity-specific cardiovascular risk prediction models across 19 ethnically diverse populations in 2.5 million individuals with COVID-19

Allery F., Pineda Moncusi M., Tomlinson C., Denaxas S., Akbar A., Denniston AK., Collins G., Delmestri A., Coates L., Bolton T., Nolan J., Lai AG., Pontikos N., Prieto Alhambra D., Sudlow C., Khunti K., Wood A., Thygesen JH., Khalid S.

Background Clinical risk stratification tools based on under-representative data risk exacerbating underlying systemic health inequalities. This study leverages a nationwide ethnically and socioeconomically diverse population for risk prediction of cardiovascular events after SARS-CoV-2 infection, by developing and evaluating prediction models driven by data from ethnicity-stratified populations. Methods This cohort study of 2,423,777 individuals uses six linked National Health Service datasets for England, UK. For each higher-level ethnicity population (Asian, Black, White, Mixed, Other, Unknown) and their 19 sub-populations, ethnicity-specific models were developed for the prediction of the risk of a cardiovascular event within one year of a SARS-CoV-2 infection. Ethnicity-specific risk prediction models were compared with a global all-ethnicity model trained on the overall study population. Models were developed using i) features used in existing tools (e.g. QRISK3), and data-driven features selected by ii) LASSO regression, iii) Random Forest, and iv) XGBoost, with corresponding SHAP analysis. Overall, 104 ethnicity-specific models using 51 candidate risk factors were developed and internally validated using discrimination (AUROC, F1, precision, recall), calibration, and net benefit analysis using a held-out test set. Findings In general, ethnicity-specific models performed similarly to the global all-ethnicity model for discrimination and net benefit, with some improvement shown in calibration. For African and Arab populations, ethnicity-specific models outperformed the global all-ethnicity model by an AUROC improvement of 2.04% (IQR:1.65%–2.25%) and 5.89% (IQR:1.95%–10.11%), respectively. Ethnicity-specific models identified additional risk factors, e.g. autoimmune liver disease in the Bangladeshi population, dementia in the African population and osteoporosis in the Gypsy and Irish Traveller population. Models using on feature selection by machine learning yielded an average AUROC improvement of 3.00% (IQR:1.10%-4.35%) compared to features informed by existing tool QRISK3. Interpretation Ethnicity-specific models based on granular data from actual practice settings have the potential to identify unique risk factors not necessarily captured in models based on dominant or majority populations, improving overall predictive performance. Our findings suggest that population-representative and ethnically diverse data can improve clinical risk stratification tailored to ethnically diverse populations and ultimately address health inequalities through targeted intervention and resource allocation.

Type

Journal article

Publisher

Elsevier

Publication Date

2026-08-20T00:00:00+00:00

Permalink More information Close