Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

INTRODUCTION: Cancer registries remain the gold standard for global cancer monitoring, yet complementing them with electronic health records and claims can significantly enhance the understanding of the cancer burden by providing a more complete picture of the patient journey. The main aim of this project is to serve as a proof of concept for using real-world data mapped to the Observational Medical Outcomes Partnership (OMOP) common data model (CDM) to monitor cancer epidemiology over time and characterise patients' clinical history and outcomes. METHODS AND ANALYSIS: This study will be conducted as an observational cohort study using a multinational network of large real-world data sources mapped to the OMOP CDM. Electronic health records (EHR) from primary and secondary care, health insurance claims and cancer registry data will be included. To date, 20 databases from 16 countries, mainly from Europe but also North America and Asia, have committed to participate in the project.We will investigate the temporal trends in incidence, prevalence and survival of 36 cancers across haematopoietic and solid tumours from 2000 (or the start of accurate data if later) to the last year with complete data. Data from all individuals registered in each of the participating data sources will be eligible for inclusion in the study. For primary care EHR and claims, individuals will be required to have at least 1 year of prior observation to ensure the identification of incident cases and adequate capture of patient characteristics. We will estimate crude and age-standardised incidence and 5-year partial prevalence. Additionally, we will estimate crude and age-standardised overall survival at 1, 5 and 10 years for the total study period and by diagnosis year groups defined according to data availability. All study objectives will be investigated at the database level, with results stratified by age and sex. For incidence and survival analyses, additional stratifications will be performed by clinical conditions and smoking status (where available). We will use the National Cancer Institute (NCI) Joinpoint Regression Programme to model overall trends in cancer incidence and the NCI JPSurv software to estimate trends in survival. Finally, we will characterise individuals diagnosed with an incident cancer based on demographics, clinical conditions and medication use at different time windows.Findings will be presented separately for each database and further summarised through descriptive aggregation by country and data source type. ETHICS AND DISSEMINATION: Each data partner will obtain study approval from their local institutional review boards prior to study execution. Distributed queries will be employed, whereby standardised analytical code is shared and run at each site locally. Deidentified, aggregated results will be returned from all participating sites. A minimum cell count of five will be used when reporting results, depending on each collaborator's data governance requirements.All study code will be publicly available, and findings will be submitted to open science journals to promote transparency and reproducibility.

More information Original publication

DOI

10.1136/bmjopen-2026-119069

Type

Journal article

Publication Date

2026-07-21T00:00:00+00:00

Volume

16

Keywords

Cancer, Epidemiology, Observational Study, Humans, Neoplasms, Incidence, Registries, Databases, Factual, Female, Electronic Health Records, Research Design, Europe, Cohort Studies, Prevalence, Male, Asia, North America