Skip to main content

Together we are beating cancer

Donate now
  • For Researchers

Data Science: unpicking data diversity and inclusive research design

The Cancer Research UK logo
by Cancer Research UK | Analysis

15 September 2026

0 comments 0 comments

Data science

Toral Gathani and Brieuc Lehmann argue that greater clarity is needed about what inclusive research design and data diversity actually mean and that scientific robustness must be maintained as research becomes more inclusive.

This entry is part 6 of 6 in the series Data science
Series Navigation<< Data science: making data count with broad consent

It’s widely recognised that a lack of diversity in biomedical datasets can undermine the fairness, accuracy and applicability of research, and contribute to inequities in scientific discovery and health outcomes.

Increasing attention is therefore being paid to the representativeness of populations included in research studies, with funders, researchers and institutions being asked to consider equality, diversity and inclusion throughout research design and delivery.

This increased focus is welcome. However, some confusion has arisen about what diversity and representativeness mean in research practice. Inclusive research can be conflated with a requirement for outcomes to be analysed according to every demographic characteristic collected. In doing so, there is a risk that analyses may be undertaken within underpowered studies leading to spurious results.

Recognition of intersectionality should not automatically result in a desire to divide and analyse datasets by increasingly smaller subgroups.

Diversity, representativeness, and intersectionality

Diversity and representativeness are often used interchangeably, but they describe different concepts.

Diversity describes the observed variation within a study population, whereas representativeness considers how well that population reflects the people to whom the research findings are intended to apply – i.e. generalisability. A study population may therefore be diverse without being fully representative.

Intersectionality adds yet another dimension to be considered. Participant characteristics such as ethnicity, age, sex, socioeconomic circumstances, disability, and geography are not independent of each other. An intersectional lens is needed to understand how these factors interact to influence health outcomes, access to healthcare, and participation in research. However, recognition of intersectionality should not automatically result in a desire to divide and analyse datasets by increasingly smaller subgroups.

Research studies should be powered to answer a specific research question. Multiple subgroup analyses reduce statistical power – that is to say, it increases the risk of false negative findings. When multiple comparisons are undertaken without appropriate statistical corrections, the possibility of apparently significant findings occurring by chance is also increased.

Both approaches may widen existing inequalities by either masking genuine differences between population groups or suggesting differences where none exist, potentially leading to inappropriate conclusions about who benefits from research and healthcare interventions.

data

Why public and patient involvement is so important

If study populations are going to be representative, consideration needs to be given as to how to achieve representativeness when studies are designed, and not after recruitment has been completed.

The importance of robust patient and public involvement at all stages of the research lifecycle is often under appreciated. The design stage needs to consider who needs to be represented, which populations may face barriers to participation and whether features of the proposed study design could inadvertently exclude particular groups of interest.

When study design is informed predominantly by people from populations already well represented in research, barriers experienced by other groups may remain unidentified and unaddressed.

The historic lack of diversity and representativeness among public contributors in research is well recognised. When study design is informed predominantly by people from populations already well represented in research, barriers experienced by other groups may remain unidentified and unaddressed.

An intersectional perspective can be valuable at this stage. Barriers to participation in research may arise through the interactions between ethnicity, age, socioeconomic circumstances, disability, geography and other factors. These issues are better considered when research is being designed than trying to do so through additional subgroup analyses after data has been collected.

Report and share but don’t necessarily analyse

There is a strong case for study populations to be reported in greater detail, regardless of whether the characteristics described are subsequently analysed, to facilitate the wider and better reuse of research data.

Consistent reporting of certain characteristics allows better metadata to be created. When it’s clear who has been included in a dataset, its relevance to other research questions can be assessed more readily and appropriate datasets can be identified for reuse. With suitable consent and governance in place, data can then be responsibly shared and, where appropriate, combined.

Populations that were too small to be reliably analysed within one study need not remain scientifically invisible. The retaining and sharing of sufficiently detailed data can facilitate individual participant level meta-analysis and other pooled approaches to allow for important research questions to be examined with greater statistical power, and at a scale at which robust answers can be obtained.

Inclusive research has rightly been prioritised, but scientific rigour should not be compromised. Sustained and informed discussions among all members of the research community about diversity, representativeness and intersectionality, the distinction between reporting and analysis, and the importance of the role of data reuse, are essential. Its only by doing this that we’ll ensure increasingly diverse datasets produce reliable and meaningful evidence, and that public investments in research are maximised.

Toral Gathani

Author

Professor Toral Gathani

Toral is Associate Professor at the Cancer Epidemiology Unit, Nuffield Department of Population Health, University of Oxford. She is also a Consultant Oncoplastic Breast Surgeon at the Oxford University Hospitals NHS Foundation Trust.

Brieuc Lehmann

Author

Professor Brieuc Lehmann

Brieuc is Associate Professor in Statistical Science at UCL

Both authors are academic co-leads of the Data Diversity Theme at Data Science for Health Equity

Data Science for Health Equity (DSxHE) is a community of practice bringing together people working in data science and health inequalities. The DSxHE Data Diversity Theme is a partnership between DSxHE and Cancer Research UK, bringing together researchers, clinicians, funders, and patient advocates to embed diversity across the research lifecycle.

Tell us what you think

Leave a Reply

Your email address will not be published. Required fields are marked *

Read our comment policy.

Tell us what you think

Leave a Reply

Your email address will not be published. Required fields are marked *

Read our comment policy.