Synthetic data to boost under-represented patients and create virtual trial cohorts: the RATE-AF case study
摘要
Clinical trials are essential for medical progress, but in certain circumstances can be constrained by recruitment costs, ethical challenges, and limited diversity in participant representation, reducing the generalisability of findings across all subgroups. In these cases, the development of digital approaches that complement traditional trials could be of particular value. We present a Bayesian network framework for synthetic data generation in clinical trials, designed to (1) boost the representation of small and under-represented subgroups and (2) generate virtual patient cohorts that replicate full trial populations with high fidelity. The framework combines probabilistic modelling with conditional synthetic data generation and is evaluated using data from the RAte control Therapy Evaluation in permanent Atrial Fibrillation (RATE-AF) randomised controlled trial, a study that compared two treatments for rate control (digoxin versus bisoprolol, a beta-blocker) in patients with atrial fibrillation and symptoms of heart failure. The framework was first assessed in a controlled boosting experiment, designed to recover simulated subgroup under-representation within the original cohort, and then extended to an exploratory boosting scenario to examine hypothetical increases in subgroup representation, before being applied to replicate the full trial population. Across these settings, it preserved statistical fidelity and reproduced the analytical results observed in the real data, boosting under-represented subgroups where sufficient data are available, whilst acknowledging limitations of boosting under extreme small-sample scenarios. This study positions synthetic data as a potential digital pathway that, with further development, could be used to support real-world clinical trials where recruitment of some population subgroups may be challenging.